A voice-over for a YouTube video
The script of a ten-minute video, read in one go. MiniMax holds the length without tiring the ear; ElevenLabs is livelier but dearer over ten minutes. The faceless YouTube video workflow takes over with the pictures.
The same script read by three speech synthesis engines, ElevenLabs v3, MiniMax Speech 2.8 HD and OpenAI TTS HD, to choose your voice-over by ear. The full workflow, no account needed.
Real workflow, real results. Pan and zoom freely.
Use this workflowThe real workflow: a Text node carries the script, three Voice-over nodes each read it with their engine, and you listen.
At a glance
A voice-over for a video, a text read aloud for a tutorial, a narration for a podcast: the question is never "can a voice be generated" but "which one". This workflow answers by ear. The script lives in a Text node and reaches three Voice-over nodes by cable, each on a different engine: change a sentence, run again, all three re-read.
The three are not there by chance. ElevenLabs v3 is the most expressive, it plays the commas and the full stops, and it is the most expensive. MiniMax Speech 2.8 HD is the most natural, with a slight accent on some proper nouns. OpenAI TTS HD is the node's default engine, steady, half the price. Each node offers its own voices, its speed, and for ElevenLabs acting tags.
The demo script runs about fifteen seconds, with a comma, a list and a hard word, enough to hear each one's prosody. The three readings cost twenty cents; the node reads text in the site's fourteen languages.
Write the script in the Text node
Careful punctuation does half the work: engines breathe at commas and drop the voice at full stops. Spell numbers out in words when the reading has to be exact.
Run the three readings
"Generate all" runs the three nodes at once. Each node shows its price before the click, computed on the length of the text.
Listen on the same passage
Compare on the hardest sentence, the one with the proper noun or the list. That is where engines part ways, not on the first sentence.
Change the voice, not the engine
Each engine has several voices, and a voice that does not fit does not condemn the engine. Try two or three on the same node before switching.
The same canvas, tuned for different needs: each variation is two or three lines of prompt away.
The script of a ten-minute video, read in one go. MiniMax holds the length without tiring the ear; ElevenLabs is livelier but dearer over ten minutes. The faceless YouTube video workflow takes over with the pictures.
Twenty seconds where every word counts: that is ElevenLabs v3's ground, which accepts acting tags in the text, a laugh, a breath, a whisper. The cinematic ad workflow shows the voice laid on the pictures.
A lesson, an article, a chapter to listen to while walking. The default engine is enough and costs the least: steadiness is a quality when you listen for an hour.
A tutorial in French, Spanish and Japanese with the same voice: ElevenLabs and MiniMax keep the timbre from one language to the next. The node reads the language of the text, you only translate the script.
Two characters answering each other: the Voice-over node offers a dialogue mode on ElevenLabs, one voice per line. The talking avatar workflow goes further, with a face that speaks.
A greeting, a shop announcement, a museum audio guide: short texts, read hundreds of times. The price per reading decides, and the default engine wins.
Every engine reads a simple sentence well. The differences show on the proper noun, the number, the list and the long sentence. That is why the demo script contains one of each.
A text written to be read reads badly aloud: sentences too long, brackets, abbreviations. Say it to yourself before giving it to the engine, and cut the sentences in two.
"$1,500 in 2024" is read three different ways depending on the engine. Write "one thousand five hundred dollars in twenty twenty-four" and all three will read the same thing.
Each engine has many voices. A voice too deep or too fast is not the engine, it is the setting: try two voices and the speed before concluding.
ElevenLabs v3 for an acted text, advertising, fiction, anything that needs intent. MiniMax for a natural, long narration, documentary or tutorial. OpenAI for volume: dozens of steady readings at a low price.
Not in this workflow, which compares the engines' voices. The Voice-over node accepts a reference voice on some engines; that is another page, and another question, the one about consent.
The site's fourteen languages for ElevenLabs and MiniMax, including Japanese, Arabic and Hebrew. OpenAI also reads most of them, with a stronger accent outside English. The node picks the language of the text.
The demo script, two hundred and thirty characters, costs seven credits on ElevenLabs, seven on MiniMax and three on OpenAI, seventeen for the three. The price follows the length of the text and shows on each node before the click.
Several thousand characters per node, the exact limit depends on the engine and shows in the node. A long text is cut into paragraphs, one node per paragraph, which also lets you re-run only the one you corrected.
The speed on all three. The emotion on MiniMax through a node setting, on ElevenLabs v3 through tags written in the text itself, in square brackets. OpenAI reads as is.
Yes, the canvas runs on mobile and the text is written directly in the node. The creating on mobile guide explains the canvas gestures by finger.
Yes, the generated voice is yours, for a public video as for commercial use. The engines' voices are licensed synthetic voices, not real people.
AI voice-over: the same text read by three engines
Use this workflowFree account, no card required. The workflow opens pre-filled in your canvas.