10% off monthly plans with code NODE10, until August 31

Imaginode
All tools

Talking AI avatar

A synthetic presenter says your script to camera, with their voice. Veo 3.1 generates image and speech together, so the lips land right. The workflow is visible here, without an account.

Interactive demo

Real workflow, real results. Pan and zoom freely.

Use this workflow

The reference sheet holds the face, the text node carries the script, the video node renders eight seconds with the voice. No voice over node: speech comes out of the video model. Open the nodes to read the prompts.

At a glance

What the workflow produces
1 image in 9:16 · 1 video of 8s in 720p, 9:16
Canvas structure
5 nodes wired by 2 connections
Models wired in
Seedream 5 Pro, Veo 3.1 Fast
Cost of one full run
about 149 credits, roughly €1.49

A talking avatar always hits the same problem: lip sync. Classic chains generate a silent video, make a voice on the side, then try to glue the two together, and there is always that tenth of a second of drift that reads as fake straight away.

Here speech and image come out of the same render. Veo 3.1 Fast generates the dialogue with the video: the lips follow the voice because they were computed together. In exchange, changing one word means running the render again, so the script is settled before you launch, which is why it lives in its own text node, editable without touching the character.

The face is held by a three angle reference sheet, the same mechanism as keeping a character consistent across several shots. That is how you produce a whole series with the same presenter, episode after episode. Seedream 5 Pro composes the opening frame, and the public calculator gives the exact cost of a take before you sign up.

How it works

  1. 1

    Create the presenter

    The reference sheet carries three views of the face. It is what lets you find exactly the same person in the next episode, in another setting and with another text.

  2. 2

    Compose the frame

    The image node frames the scene: office, light, room left for on screen text. A chest up shot slightly off centre leaves space for hand gestures.

  3. 3

    Write the line

    The text node holds the staging and the sentence to say, in quotes. Two short sentences are enough for eight seconds, and the tone is described as a direction to an actor.

  4. 4

    Render the take

    Eight seconds, vertical, audio on. The voice arrives inside the video file: there is nothing to sync afterwards, and nothing to edit if the take is good.

Variations and use cases

The same canvas, tuned for different needs: each variation is two or three lines of prompt away.

Presenter avatar for a training course

A course splits into eight-second segments, each in its own Video node, all wired to the same face card. The presenter stays identical from chapter to chapter, which no series of real takes guarantees at this price.

Brand spokesperson

The same face opens every company video, without depending on a colleague's availability or their leaving. Describe the outfit and the set in the Image node: the card holds the face, the shot holds the context.

Avatar speaking another language

Write the line in quotes in the language you want: Veo 3.1 Fast pronounces it, lips included. One video then ships to five markets, with the same presenter and five lines.

Welcome video on a landing page

Eight seconds, chest-up shot, blurred office background: the format that lifts time on page. Keep the sentence short, one idea per video, and plan a subtitle for sound-off viewing.

Vertical avatar for TikTok and Reels

The template already renders 9:16. Leave headroom for on-screen text, and cut the first half second at edit time: video models often start on a slightly frozen frame.

Common mistakes

A script too long for eight seconds

Two short sentences fit in one take. Beyond that the model speeds up or cuts the ending, and fixing one word means paying for the whole render again.

Describing the voice after launching the render

The voice comes out of the same render as the image: timbre, pace and intent are described in the prompt, before generating. Changing the voice means a new take.

A tight close-up on the face

Very tight shots expose the micro flaws around the mouth. A chest-up shot, slightly off centre, leaves room for hand gestures and reads as more credible.

Frequently asked questions

Can the voice be chosen?

You describe it in the prompt, like a casting note: timbre, pace, intent. It is not picked from a list, and two renders of the same text can differ slightly.

Does the model speak languages other than English?

Yes, and the template does it in French. Write the line in quotes in the language you want, the rest of the description can stay in English without changing the spoken language.

What if the take is bad?

You run it again. An eight second take costs a few tens of cents of credits, and the node keeps the history of previous renders: nothing is lost when you compare two versions.

Is the lip sync accurate?

Yes, because image and speech come out of the same render: nothing is stitched afterwards. That is what separates this workflow from classic chains, where a tenth of a second of drift gives the edit away.

Can I use my own face?

Yes, by filling the Reference card with three photos of yourself. Only do this with a face whose owner agreed: a spokesperson generated from someone else's photos is a real image-rights problem.

What does an eight-second take cost?

About 144 credits with Veo 3.1 Fast, roughly 1.40 euros, opening image included. The exact price shows on the button before you launch the render.

Can I chain several lines?

Yes, one Video node per line, all wired to the same face card and the same opening shot. Cut end to end, that gives a one-minute video with a consistent presenter.

Talking AI avatar

Use this workflow

Free account, no card required. The workflow opens pre-filled in your canvas.