All tools

Add sound to a video with AI

A silent shot, an ambience bed and a music bed: the three pieces are generated here, the montage assembles them. The full workflow, no account needed.

Interactive demo

Real workflow, real results. Pan and zoom freely.

Use this workflow

The real workflow: the rainy terrace shot is silent, and the two nodes produce the ambience and the music it was missing.

At a glance

What the workflow produces
2 audio tracks, 40 s in total
Canvas structure
7 nodes wired by 5 connections
Models wired in
ElevenLabs Sound Effects, ElevenLabs Music
Cost of one full run
about 87 credits, roughly $1.04

Most generated videos come out silent, and that silence gives them away before anything else does. This workflow builds the two missing tracks: an ambience bed with ElevenLabs Sound Effects, a music bed with ElevenLabs Music.

The Media node is not wired to anything, and that is not an oversight. No model in the catalogue takes a video as input to derive its sound effects: sound is described in text. So the shot stays in front of you while you write the two instructions, which is exactly what it is there for.

Two separate nodes rather than one, because an ambience and a music bed are not set the same way: the first is matched to the length of the shot, the second is generated longer and cut in the edit. The assembly happens in the Montage, the tool in the canvas toolbar, which exports a single file. A full run costs just over a dollar.

How it works

  1. 1

    Import your silent shot

    A generated video, a phone rush, a screen capture: the Media node takes it as is. Note its length, that is what decides how long the ambience should be.

  2. 2

    Describe what should be heard

    Not the emotion, the sources. Rain on canvas, water dripping into a puddle, tyres on wet tarmac. Sound effects are described like an inventory, not like an intention.

  3. 3

    Describe the music separately

    Instruments, tempo, no drums, no vocals. Say it must leave room for the effects: music that fills the whole spectrum crushes the ambience you have just generated.

  4. 4

    Assemble in the Montage

    Video, ambience and music on three tracks, levels set, and one file out. That is where the shot stops being silent.

Variations and use cases

The same canvas, tuned for different needs: each variation is two or three lines of prompt away.

Adding sound to a generated shot

The direct use case: every video out of a Video node without audio, which is most of them. Take the description that generated the image, keep the sources and drop the intentions, and you already have three quarters of the ambience instruction.

An ambience loop for a site or a kiosk

A shot looping at the top of a page is better off carrying only the ambience, no music. The effects are set to the exact length of the loop so the join cannot be heard, which cut music never allows.

Replacing a disappointing original sound

A phone rush brings back wind in the microphone and the neighbour's voice. Generating a clean ambience and swapping it in the edit costs less than a reshoot, and is less noticeable than a silent shot.

Music alone for an edit

The Music node works without the rest of the workflow: thirty seconds of instrumental bed for a cut of several shots. Generate longer than you need, the cut happens in the edit and a fade beats an abrupt ending.

One-off effects rather than a bed

A door, a glass set down, a footstep on gravel. ElevenLabs Sound Effects can be short and precise, and those one-off sounds are placed to the frame in the edit. One instruction per sound, never a list.

Nothing forces you to start from a video. A run of stills cut as a slideshow is scored exactly the same way, and it is often what makes a carousel pass for a real sequence.

Common mistakes

Describing an emotion instead of a source

A melancholic, heavy ambience means nothing to a sound effects engine. Rain on canvas, water dripping, tyres on wet tarmac: that is what it can do. The emotion is the result, not the instruction.

Asking for music and effects together

The mistake that gives the most disappointing result. The model blends the two registers and returns a hybrid where the rain sounds like a shaker. Two nodes, two instructions, two levels you can set in the edit.

Forgetting to ban vocals

A music model left free will happily add singing, and singing under a video rules out any voice-over afterwards. Write no vocals in the instruction, even when it seems obvious.

Generating the ambience at the length of the edit

The ambience matches the shot, not the film. Twenty seconds of rain for a five-second shot forces an arbitrary cut, whereas a bed generated at the right length drops in without a thought.

Frequently asked questions

Why is the Media node not connected to anything?

Because no model knows how to listen to a picture yet. Effects and music are generated from text, not from the video. The shot is there to be looked at while you write, and then to go to the montage with the two tracks.

Do I really need two nodes?

Yes, and it is the most useful point of the template. A single instruction asking for rain and a piano gives you mush where the effects sound like an instrument. Separated, each is set to its own length and level.

Can I generate a voice-over here?

Not with these two nodes, which do effects and music. Voice has its own node, and the documentary video workflow shows how it wires to a text.

What does a full run cost?

Just over a dollar for the ten seconds of ambience and the thirty seconds of music. The price shows on each node before you click, and it depends on the length you ask for.

Does the montage really export a single file?

Yes, video and audio tracks mixed into an MP4, encoded in the browser. It is the Montage tool in the canvas toolbar, which takes the project's renders without you having to download them one by one.

Can I use the music commercially?

Yes, the renders are yours. The caution is on what you ask for: an instruction naming an artist or an existing track is chasing imitation, and that is where the ground gets slippery.

What is the maximum length?

Twenty-two seconds for a sound effect, several minutes for music. For a longer ambience, generating a loopable bed and repeating it in the edit beats asking for a very long one.

How is this different from sound generated by the video model?

Some video models produce their own soundtrack, which is handy when it lands right. Here you keep control: two separate tracks, adjustable, replaceable, and the shot stays usable even if the ambience does not fit.

Add sound to a video with AI

Use this workflow

Free account, no card required. The workflow opens pre-filled in your canvas.