All tools

AI voice-over: the same text read by three engines

The same script read by three speech synthesis engines, ElevenLabs v3, MiniMax Speech 2.8 HD and OpenAI TTS HD, to choose your voice-over by ear. The full workflow, no account needed.

Interactive demo

Real workflow, real results. Pan and zoom freely.

Use this workflow

The real workflow: a Text node carries the script, three Voice-over nodes each read it with their engine, and you listen.

At a glance

What the workflow produces
3 voice-overs
Canvas structure
5 nodes wired by 3 connections
Models wired in
ElevenLabs v3, MiniMax Speech 2.8 HD, OpenAI TTS HD
Cost of one full run
about 17 credits, roughly $0.20

A voice-over for a video, a text read aloud for a tutorial, a narration for a podcast: the question is never "can a voice be generated" but "which one". This workflow answers by ear. The script lives in a Text node and reaches three Voice-over nodes by cable, each on a different engine: change a sentence, run again, all three re-read.

The three are not there by chance. ElevenLabs v3 is the most expressive, it plays the commas and the full stops, and it is the most expensive. MiniMax Speech 2.8 HD is the most natural, with a slight accent on some proper nouns. OpenAI TTS HD is the node's default engine, steady, half the price. Each node offers its own voices, its speed, and for ElevenLabs acting tags.

The demo script runs about fifteen seconds, with a comma, a list and a hard word, enough to hear each one's prosody. The three readings cost twenty cents; the node reads text in the site's fourteen languages.

How it works

  1. 1

    Write the script in the Text node

    Careful punctuation does half the work: engines breathe at commas and drop the voice at full stops. Spell numbers out in words when the reading has to be exact.

  2. 2

    Run the three readings

    "Generate all" runs the three nodes at once. Each node shows its price before the click, computed on the length of the text.

  3. 3

    Listen on the same passage

    Compare on the hardest sentence, the one with the proper noun or the list. That is where engines part ways, not on the first sentence.

  4. 4

    Change the voice, not the engine

    Each engine has several voices, and a voice that does not fit does not condemn the engine. Try two or three on the same node before switching.

Variations and use cases

The same canvas, tuned for different needs: each variation is two or three lines of prompt away.

A voice-over for a YouTube video

The script of a ten-minute video, read in one go. MiniMax holds the length without tiring the ear; ElevenLabs is livelier but dearer over ten minutes. The faceless YouTube video workflow takes over with the pictures.

A twenty-second advert

Twenty seconds where every word counts: that is ElevenLabs v3's ground, which accepts acting tags in the text, a laugh, a breath, a whisper. The cinematic ad workflow shows the voice laid on the pictures.

A text read aloud for learning

A lesson, an article, a chapter to listen to while walking. The default engine is enough and costs the least: steadiness is a quality when you listen for an hour.

The same voice in several languages

A tutorial in French, Spanish and Japanese with the same voice: ElevenLabs and MiniMax keep the timbre from one language to the next. The node reads the language of the text, you only translate the script.

A two-voice dialogue

Two characters answering each other: the Voice-over node offers a dialogue mode on ElevenLabs, one voice per line. The talking avatar workflow goes further, with a face that speaks.

A voice for an answering machine or a guide

A greeting, a shop announcement, a museum audio guide: short texts, read hundreds of times. The price per reading decides, and the default engine wins.

Common mistakes

Judging on the first sentence

Every engine reads a simple sentence well. The differences show on the proper noun, the number, the list and the long sentence. That is why the demo script contains one of each.

Writing for the eye, not the ear

A text written to be read reads badly aloud: sentences too long, brackets, abbreviations. Say it to yourself before giving it to the engine, and cut the sentences in two.

Leaving numbers as digits

"$1,500 in 2024" is read three different ways depending on the engine. Write "one thousand five hundred dollars in twenty twenty-four" and all three will read the same thing.

Switching engines before switching voices

Each engine has many voices. A voice too deep or too fast is not the engine, it is the setting: try two voices and the speed before concluding.

Frequently asked questions

Which of the three should I pick?

ElevenLabs v3 for an acted text, advertising, fiction, anything that needs intent. MiniMax for a natural, long narration, documentary or tutorial. OpenAI for volume: dozens of steady readings at a low price.

Can I use my own voice?

Not in this workflow, which compares the engines' voices. The Voice-over node accepts a reference voice on some engines; that is another page, and another question, the one about consent.

Which languages does it work in?

The site's fourteen languages for ElevenLabs and MiniMax, including Japanese, Arabic and Hebrew. OpenAI also reads most of them, with a stronger accent outside English. The node picks the language of the text.

How much does a voice-over cost?

The demo script, two hundred and thirty characters, costs seven credits on ElevenLabs, seven on MiniMax and three on OpenAI, seventeen for the three. The price follows the length of the text and shows on each node before the click.

What is the maximum text length?

Several thousand characters per node, the exact limit depends on the engine and shows in the node. A long text is cut into paragraphs, one node per paragraph, which also lets you re-run only the one you corrected.

Can I set the emotion or the speed?

The speed on all three. The emotion on MiniMax through a node setting, on ElevenLabs v3 through tags written in the text itself, in square brackets. OpenAI reads as is.

Can I do it from my phone?

Yes, the canvas runs on mobile and the text is written directly in the node. The creating on mobile guide explains the canvas gestures by finger.

Is the file free to use?

Yes, the generated voice is yours, for a public video as for commercial use. The engines' voices are licensed synthetic voices, not real people.

AI voice-over: the same text read by three engines

Use this workflow

Free account, no card required. The workflow opens pre-filled in your canvas.