A phone interview
The textbook case: two people in a café, the phone set between them. The filter removes the cups and the coffee machine, both voices stay, with the distance that comes with them. One pass is usually enough.
A voice memo dictated on a noisy terrace, run through both noise removers of the catalogue side by side, DeepFilterNet 3 and ElevenLabs Audio Isolation. The full workflow, no account needed.
Real workflow, real results. Pan and zoom freely.
Use this workflowThe real workflow: the same noisy memo enters two Enhance audio nodes, and you hear what each one removed.
At a glance
An interview recorded on a phone in a café, a voice memo dictated in the street, a podcast voice with the fridge behind it: background noise is what separates a usable recording from one to redo. This workflow removes it with the two catalogue models that know how, and puts them side by side because they do not work the same way.
DeepFilterNet 3 is a filter: it removes what is not voice and leaves the air of the room, the result stays natural. ElevenLabs Audio Isolation rebuilds the voice: drier, more radio-like, sometimes a little synthetic on consonants. On the demo memo, built on purpose by mixing a clean voice with a street ambience and a hum, we know exactly what there was to remove.
Your file enters through the node's import button, up to ten minutes per run, and the price is counted by the second: both clean-ups on fourteen seconds cost seven cents, and DeepFilterNet alone comes to nine credits per minute.
Import the recording
MP3, WAV, M4A or OGG, up to ten minutes. The node reads the duration and shows the price before any click. An audio produced elsewhere on the canvas can also be cabled to the port.
Run the filter
DeepFilterNet 3 first: it is the cheapest and the most faithful. If the voice comes out clear and the room is still there, you are done.
Try the rebuild
If the noise was loud, or if the voice has to sit on music, run isolation next to it. It comes out drier, which mixes better in an edit.
Listen to the consonants
That is where the two part ways: the s and t sounds are what a noise remover damages first. The winning file goes to the edit or downloads from the node.
The same canvas, tuned for different needs: each variation is two or three lines of prompt away.
The textbook case: two people in a café, the phone set between them. The filter removes the cups and the coffee machine, both voices stay, with the distance that comes with them. One pass is usually enough.
The fridge, the street through the window, the computer fan: a constant, soft noise that the filter removes without a trace. On a voice close to the mic, DeepFilterNet is what you need, the rebuild would be too much.
The demo memo. Loud, variable noise, cars and passers-by. The filter leaves a bed, the rebuild erases it: here it wins, because you want to understand, not keep the atmosphere.
A voice-over filmed without a studio, to lay on pictures and music. The isolation result, drier, blends better in the Montage than a voice still dragging the room behind it.
A digitised tape, with its hiss and its hum. The filter removes the hiss and keeps the voice as it was, which is what you want from a memory: clear, not rewritten.
A distant voice, reverb, coughs. Both clean the noise; neither shortens the room's reverb, which is not noise but space. Expect a clear voice that is still far away.
The price is counted by the second. Cleaning an hour of recording to keep ten minutes costs six times too much. Cut first, clean next, on the passage you keep.
The filter then the rebuild, or the reverse, does not give a twice-better clean-up: the second model works on an already transformed voice and amplifies its artefacts. One pass, chosen by ear.
A noise remover removes what is not the voice; it does not bring a mic that was three metres away any closer. Reverb and distance remain. What you will hear is the same voice, without the noise around it.
A concert recording run through a noise remover comes out with the voice and without the band: the instruments are treated as noise. For a song, it is stem separation you need, not clean-up.
The filter first, almost always: it removes the noise without rewriting the voice. The rebuild when the noise was louder than the voice, or when the result has to blend into music, because a drier timbre mixes better.
No, and that is by design. Both models keep the voice and throw away the rest: on a song, the music is treated as noise. To split a song into stems, the split the vocals from a song workflow does that job with Demucs.
Both voices are kept, neither model tells speakers apart. What leaves is the noise between them, not one of them. To keep a single voice, you describe what you want to SAM Audio, another node of the catalogue.
DeepFilterNet 3 comes to nine credits per minute, ElevenLabs isolation to fourteen. The demo memo, fourteen seconds in both nodes, costs six credits. The price shows on the node before you click, on the measured duration of your file.
Ten minutes per run. A longer recording is cut into pieces in the edit, or cleaned only on the passages you keep, which is cheaper anyway.
The filter, no: it removes, it adds nothing. The rebuild, a little: it rewrites the voice from what it recognised, and on sibilant consonants you can hear it. That is why the demo shows both.
Yes, in two steps: the Montage exports the soundtrack, the node cleans it, and the clean voice comes back to sit on the pictures in place of the old one.
No. What you upload stays in your workspace and is handed to no training, and the question matters more here than elsewhere: we are often talking about interviews and the voices of relatives.
Clean up a recording with AI
Use this workflowFree account, no card required. The workflow opens pre-filled in your canvas.