A company podcast without a studio
One script per episode, two fixed voices, a minute per node. The edit assembles the nodes, adds a jingle and music. No microphone, no room, no bad take.
An eight-line exchange between two hosts, performed in a single take by ElevenLabs Dialogue v3: two voices, the hand-offs, the laughs. The full workflow, visible without an account, with the episode.
Real workflow, real results. Pan and zoom freely.
Use this workflowReal workflow: a single Voice-over node in dialogue mode, eight lines, one voice per line, on ElevenLabs Dialogue v3. Listen to the episode.
At a glance
A two-voice podcast is not made by gluing two voice-overs one after the other: the silences between lines fall flat, the pick-ups are too clean, nobody laughs at the right moment. ElevenLabs Dialogue v3 receives all the lines together and performs the exchange in one take, like two people in the same room. The demo lasts a minute, on a simple question: why time seems to go faster as we get older.
Everything fits in one Voice-over node in dialogue mode. Each line has its own voice, picked from the model's twenty-one voices; the demo takes one female and one male, because two similar voices sound like a monologue. Bracketed tags, [laughs], [curious], [whispers], direct the acting without being read aloud: the fifth line of the demo starts with [laughs], and you hear it.
Ten lines at most per node, three thousand characters. For a longer episode, chain several nodes and assemble them in the Edit node. The model is billed per character: count on about a dozen credits per minute of speech. An hour of podcast comes to under twelve dollars.
Write the exchange, line by line
Short sentences, the way people talk. One line per idea. Have them answer, push back, interrupt: the back and forth makes the podcast, not the length of the speeches.
Give each line a voice
One female, one male, far apart in timbre. The same voice for the same host from start to finish. The node offers the list, listen to two before choosing.
Place the acting tags
[laughs] before a line that is amused, [curious] before a question, [sighs] before a concession. One tag per line at most, the model does the rest.
Generate, listen, fix one line
The node performs the whole exchange. If a line sounds wrong, change its text or its tag and rerun: the model replays everything, consistent from end to end.
The same canvas, tuned for different needs: each variation is two or three lines of prompt away.
One script per episode, two fixed voices, a minute per node. The edit assembles the nodes, adds a jingle and music. No microphone, no room, no bad take.
Have one voice ask the questions and the other answer: a blog post becomes a five-minute exchange you listen to while walking. The canvas text model can write the script.
Teacher and student, customer and adviser, patient and doctor. Two voices, hesitation tags, and the case comes alive. Each scenario in its own node.
Two characters, emotion tags, [whispers], [sad], [excited], and a generated sound effect under the scene. The Edit node lays the voices on the ambience.
Thirty seconds, a question and an answer, the brand in the last line. The female and the male voice share the roles, the edit adds the music.
The file comes out clean; if a level needs smoothing, the Enhance node takes care of it. The MP3 downloads from the node, ready for the podcast platform.
Two female voices of the same age, or two deep male ones: you no longer know who is speaking. Take one voice of each timbre, as the demo does, and keep them from one episode to the next.
A two-hundred-word line is no longer an exchange, it is a lecture. Cut it, have the other voice pick it up. The pace of a podcast comes from the back and forth.
One tag per line directs the acting; four make a caricature. [laughs] on every line ends up sounding fake. The demo uses two across eight lines.
The node accepts ten, and three thousand characters, so the model keeps the exchange consistent. Beyond that, split into several nodes and assemble them in the edit.
Yes. The model receives the lines together and handles the hand-offs, the breaths and the pace of the exchange. That is the difference with two voice-overs joined in the edit.
Not here. The model has twenty-one voices, it does not clone. To read a text with a voice that resembles yours, the change your voice workflow replays your recording with another timbre.
The language of your lines, in the site's fourteen languages. The language is set on the node; the demo is in English, with voices that pronounce it.
The price is computed per character and shows on the node before you click, about a dozen credits per minute of speech. The demo exchange, a thousand characters, costs fourteen. An hour of podcast comes to under twelve dollars.
Twenty-one, female and male, all able to speak the site's fourteen languages. A dialogue can mix more than two, but two are enough for a podcast.
Three thousand characters, about three minutes of speech. A twenty-minute episode is seven nodes chained in the Edit node, each one rerunnable on its own.
Yes, in the same canvas: a Music node for the theme, and the Edit node to lay the voices over it. The add sound to a video workflow shows the principle.
Yes. The canvas runs on mobile, the lines are typed in the node and the episode plays inside it. The create on mobile guide explains the gestures.
Make a two-voice podcast with AI
Use this workflowFree account, no card required. The workflow opens pre-filled in your canvas.