All tools

Split the vocals from a song with AI

A 30-second song split into four stems by Demucs, vocals, drums, bass and the rest, with the same song run through vocal isolation next to it. The full workflow, no account needed.

Interactive demo

Real workflow, real results. Pan and zoom freely.

Use this workflow

The real workflow: the same song enters two Enhance audio nodes, Demucs returns four stems, isolation returns one.

At a glance

What the workflow produces
5 processed audio files
Canvas structure
3 nodes wired by 0 connections
Models wired in
Demucs, ElevenLabs Audio Isolation
Cost of one full run
about 12 credits, roughly $0.14

Making a karaoke, isolating the voice for a remix, pulling the drums out of a track to play along: it all starts with the same operation, splitting a song into its stems. This workflow does it with Demucs, Meta's separation model: the song enters an Enhance audio node and four files come out, the vocals, the drums, the bass and everything else, guitars and keys included.

The second node deliberately does something else. ElevenLabs Audio Isolation is built for speech, not music, and the demo shows what it does with a song anyway: it returns the voice alone. Measured on 3 September 2026 on the demo song, its envelope follows the Demucs vocal stem at 0.93 correlation, 17% louder. But that is all it returns: no instrumental, so no karaoke.

The demo song is generated, so there are no rights to ask anyone for. Yours enters through each node's import button, up to ten minutes per run, and the price is counted by the second: both nodes on thirty seconds cost fourteen cents, four credits of which go to Demucs.

How it works

  1. 1

    Import your song

    MP3, WAV, M4A or OGG, up to ten minutes. The node shows the duration as soon as the player has read it, and the price with it. The audio port also accepts music generated on the canvas.

  2. 2

    Run Demucs

    One click, four files: vocals, drums, bass, other. Each stem has its own player and download button, and the node's output carries the first one, the vocals.

  3. 3

    Compare with isolation

    The second node gives the voice alone by another route. On a heavily produced vocal, one or the other keeps the line endings better: listen to both before choosing which one goes to the edit.

  4. 4

    Assemble your version

    The karaoke is the three stems without the voice: lay them on three audio tracks of the [Montage](/docs/canvas-nodes) and export. For a remix, keep the voice and replace the rest.

Variations and use cases

The same canvas, tuned for different needs: each variation is two or three lines of prompt away.

A karaoke version

The most requested case: the song without its voice. Demucs puts the vocals aside, the karaoke is what remains, drums, bass and other. Three stems to lay in the edit, and a song to sing over.

An a cappella for a remix

The reverse: the voice alone, to lay on a new production. Here the two nodes compete, and that is the point of the demo: Demucs and isolation each return a slightly different voice, take the one whose line endings hold best.

Playing along to a track's drums

A drummer learning a track isolates the drum stem and loops it, without the guitar on top. The bass works the same way. These are the two stems Demucs separates most cleanly.

Cleaning up a cover recorded on a phone

A cover filmed in a living room has the voice and the guitar mixed together. Extract the audio from the video, run it through Demucs, and the voice can be remixed separately, louder, before going back on the pictures.

Isolating a single instrument

Demucs only knows four families. To pull out precisely a guitar or a piano, the SAM Audio node takes a text description, "the electric guitar", and returns what it recognised on one side, the rest on the other.

Subtitling or transcribing a song

A separated voice transcribes far better than a mix: the words no longer fight the snare. Separate first, transcribe next, and keep the vocal stem to check a doubtful word.

Common mistakes

Asking vocal isolation for a karaoke

It returns the voice, never its opposite. The node says "isolation", and that is exactly what it does: what is not voice is thrown away, not set aside. For the karaoke, it is Demucs.

Splitting an already heavily compressed file

A 96 kb/s MP3 or a song recorded from a speaker has already lost the highs Demucs uses to recognise instruments. The stems come out blurry. Start from the best file available, the original if possible.

Running ten minutes to keep thirty seconds

The price is counted by the second, and Demucs is slow on long files. Cut to the passage you care about first, then split: cheaper, faster, and the result is the same.

Expecting a "guitar" stem

Demucs puts everything that is neither voice, drums nor bass into "other": guitars, keys, strings, together. For a specific instrument, it is the SAM Audio node described above, not a fifth stem that does not exist.

Frequently asked questions

Which of the two nodes makes the karaoke?

Demucs, and only Demucs. The karaoke is everything that is not the voice, that is the drums, bass and other stems together. Vocal isolation only returns the voice, so it cannot produce the opposite.

Is the separated voice clean?

Clean to the ear, not perfect under a magnifying glass. On backing vocals and long reverbs, Demucs sometimes leaves a breath of the rest of the mix in the vocal stem. In a remix it drowns under the new instruments; for a bare a cappella, listen to the silences.

Can I split a copyrighted track?

Technically yes, the model does not read rights. What you do with the result falls under the track's licence: a karaoke for home raises no question, a publication does. The demo song is generated precisely for that reason.

How much does a separation cost?

Four credits for thirty seconds with Demucs, six per minute, about sixty for the maximum ten minutes. Vocal isolation is dearer, fourteen credits per minute. The price shows on the node before you click, on the measured duration of your file.

What is the maximum duration?

Ten minutes per run, for both nodes. A whole album is processed track by track, which is what you want anyway: one stem set per song.

What format comes out?

Audio files readable in any software, one per stem, downloadable from the node. Each stem has the exact duration of the original song, they line up without any offset.

Can I do it from my phone?

Yes, the canvas runs on mobile and the file is chosen from the node, including from the phone's files. The creating on mobile guide explains the canvas gestures by finger.

Are my files used to train models?

No. What you upload stays in your workspace and is handed to no training. The produced files are yours, within the rights of the original track.

Split the vocals from a song with AI

Use this workflow

Free account, no card required. The workflow opens pre-filled in your canvas.