All tools

Clean up a recording with AI

A voice memo dictated on a noisy terrace, run through both noise removers of the catalogue side by side, DeepFilterNet 3 and ElevenLabs Audio Isolation. The full workflow, no account needed.

Interactive demo

Real workflow, real results. Pan and zoom freely.

Use this workflow

The real workflow: the same noisy memo enters two Enhance audio nodes, and you hear what each one removed.

At a glance

What the workflow produces
2 processed audio files
Canvas structure
3 nodes wired by 0 connections
Models wired in
DeepFilterNet 3, ElevenLabs Audio Isolation
Cost of one full run
about 6 credits, roughly $0.07

An interview recorded on a phone in a café, a voice memo dictated in the street, a podcast voice with the fridge behind it: background noise is what separates a usable recording from one to redo. This workflow removes it with the two catalogue models that know how, and puts them side by side because they do not work the same way.

DeepFilterNet 3 is a filter: it removes what is not voice and leaves the air of the room, the result stays natural. ElevenLabs Audio Isolation rebuilds the voice: drier, more radio-like, sometimes a little synthetic on consonants. On the demo memo, built on purpose by mixing a clean voice with a street ambience and a hum, we know exactly what there was to remove.

Your file enters through the node's import button, up to ten minutes per run, and the price is counted by the second: both clean-ups on fourteen seconds cost seven cents, and DeepFilterNet alone comes to nine credits per minute.

How it works

  1. 1

    Import the recording

    MP3, WAV, M4A or OGG, up to ten minutes. The node reads the duration and shows the price before any click. An audio produced elsewhere on the canvas can also be cabled to the port.

  2. 2

    Run the filter

    DeepFilterNet 3 first: it is the cheapest and the most faithful. If the voice comes out clear and the room is still there, you are done.

  3. 3

    Try the rebuild

    If the noise was loud, or if the voice has to sit on music, run isolation next to it. It comes out drier, which mixes better in an edit.

  4. 4

    Listen to the consonants

    That is where the two part ways: the s and t sounds are what a noise remover damages first. The winning file goes to the edit or downloads from the node.

Variations and use cases

The same canvas, tuned for different needs: each variation is two or three lines of prompt away.

A phone interview

The textbook case: two people in a café, the phone set between them. The filter removes the cups and the coffee machine, both voices stay, with the distance that comes with them. One pass is usually enough.

A podcast voice recorded at home

The fridge, the street through the window, the computer fan: a constant, soft noise that the filter removes without a trace. On a voice close to the mic, DeepFilterNet is what you need, the rebuild would be too much.

A memo dictated in the street

The demo memo. Loud, variable noise, cars and passers-by. The filter leaves a bed, the rebuild erases it: here it wins, because you want to understand, not keep the atmosphere.

A voice going into a video

A voice-over filmed without a studio, to lay on pictures and music. The isolation result, drier, blends better in the Montage than a voice still dragging the room behind it.

An old family recording

A digitised tape, with its hiss and its hum. The filter removes the hiss and keeps the voice as it was, which is what you want from a memory: clear, not rewritten.

A talk recorded from the audience

A distant voice, reverb, coughs. Both clean the noise; neither shortens the room's reverb, which is not noise but space. Expect a clear voice that is still far away.

Common mistakes

Cleaning before cutting

The price is counted by the second. Cleaning an hour of recording to keep ten minutes costs six times too much. Cut first, clean next, on the passage you keep.

Stacking both treatments

The filter then the rebuild, or the reverse, does not give a twice-better clean-up: the second model works on an already transformed voice and amplifies its artefacts. One pass, chosen by ear.

Expecting a studio from a distant recording

A noise remover removes what is not the voice; it does not bring a mic that was three metres away any closer. Reverb and distance remain. What you will hear is the same voice, without the noise around it.

Cleaning music by mistake

A concert recording run through a noise remover comes out with the voice and without the band: the instruments are treated as noise. For a song, it is stem separation you need, not clean-up.

Frequently asked questions

Which of the two should I choose?

The filter first, almost always: it removes the noise without rewriting the voice. The rebuild when the noise was louder than the voice, or when the result has to blend into music, because a drier timbre mixes better.

Does it work on music?

No, and that is by design. Both models keep the voice and throw away the rest: on a song, the music is treated as noise. To split a song into stems, the split the vocals from a song workflow does that job with Demucs.

What if two people are talking?

Both voices are kept, neither model tells speakers apart. What leaves is the noise between them, not one of them. To keep a single voice, you describe what you want to SAM Audio, another node of the catalogue.

How much does a clean-up cost?

DeepFilterNet 3 comes to nine credits per minute, ElevenLabs isolation to fourteen. The demo memo, fourteen seconds in both nodes, costs six credits. The price shows on the node before you click, on the measured duration of your file.

What is the maximum duration?

Ten minutes per run. A longer recording is cut into pieces in the edit, or cleaned only on the passages you keep, which is cheaper anyway.

Does the clean-up change the voice?

The filter, no: it removes, it adds nothing. The rebuild, a little: it rewrites the voice from what it recognised, and on sibilant consonants you can hear it. That is why the demo shows both.

Can I clean the audio of a video?

Yes, in two steps: the Montage exports the soundtrack, the node cleans it, and the clean voice comes back to sit on the pictures in place of the old one.

Are my recordings used to train models?

No. What you upload stays in your workspace and is handed to no training, and the question matters more here than elsewhere: we are often talking about interviews and the voices of relatives.

Clean up a recording with AI

Use this workflow

Free account, no card required. The workflow opens pre-filled in your canvas.