All tools

Make a photo talk with AI

A face photo and a text read by a catalog voice: InfiniTalk makes the person speak, lips, blinks and head movements included. The video lasts as long as the voice. The whole workflow is visible without an account.

Interactive demo

Real workflow, real results. Pan and zoom freely.

Use this workflow

The real workflow: a photo in a Media node, two texts each read by a Voice-over node, and two Video nodes set to InfiniTalk that receive the photo through the image port and the voice through the voice port.

At a glance

What the workflow produces
2 videos of 18s in 480p, 1:1 · 2 voice-overs
Canvas structure
8 nodes wired by 6 connections
Models wired in
ElevenLabs v3, InfiniTalk
Cost of one full run
about 446 credits, roughly $5.35

A single photo is enough. No video to shoot, no avatar to configure: the face in the photo starts talking, with the lips timed to every syllable, the blinks, the small head movements that go with a sentence. The result watches like a to-camera take, and it comes from a phone portrait.

The text is read in a Voice-over node, with one of the catalog voices (ElevenLabs v3 first) or a cloned voice. The Video node set to InfiniTalk receives the photo and the voice, and renders a video that lasts exactly as long as the audio, up to twenty-eight seconds. The price follows that duration, measured on the file, and shows before you click.

It is the gesture of the talking avatar, but from YOUR face, or that of a generated character, a drawn mascot, an old portrait. A birthday message, a product presentation, an announcement on social networks: the demo shows two messages on the same photo.

How it works

  1. 1

    Pick a photo facing the camera

    Sharp face, mouth closed and visible, front light, plain background. A phone portrait works very well. Glasses and beards pass, a hand in front of the mouth or a profile does not.

  2. 2

    Write the text and pick the voice

    The Voice-over node reads what you write with the chosen voice, in your language. A twenty-word sentence makes seven to eight seconds. Generate the voice first: it sets the duration and the price of the video.

  3. 3

    Generate the video

    The InfiniTalk node shows the measured duration of the voice and the exact price. Click: the face talks, at 480p for social networks or 720p for a render to project.

  4. 4

    Chain if the text is long

    Beyond twenty-eight seconds of voice, split the text into two takes and edit them back to back in a Montage node, for free, as in the other video workflows.

Variations and use cases

The same canvas, tuned for different needs: each variation is two or three lines of prompt away.

The birthday message

The photo of the person being celebrated, a twenty-word text read by a warm voice: that is the second example of the demo. The video is sent as is by message, and nobody had to film themselves.

The product presentation without a shoot

A generated portrait becomes a presenter: the launch text in the Voice-over node, the photo in the Media node, and the video comes out at 720p, ready for a product page or a UGC ad.

An old portrait that tells its story

A scanned family photo, a text written in the first person: the face comes alive and speaks. Run the photo through an Enhance node first if it is small or damaged, then through InfiniTalk.

The mascot that announces

A drawn character with a readable mouth works: the model animates the lips of a drawing like those of a photo. Ideal for a brand announcement on social networks, with a cheerful catalog voice.

The same message in three languages

One photo, three Voice-over nodes in three languages, three InfiniTalk nodes: the same person speaks English, Spanish and French, with accurate lips every time. To start from a video already shot, see dubbing.

The photo that talks, then moves

Take the produced video and run it through make a photo dance or into a Video node as a reference: speech first, body next.

Common mistakes

A face in profile or a tight three-quarter

The model needs to see the whole mouth. Frontal or in a slight three-quarter, the sync is clean; in profile, the lips distort.

An open or hidden mouth on the photo

A smile with visible teeth, a hand on the chin, a microphone in front: so many obstacles. Mouth closed, face clear, and the result is sharp.

Writing the text in the Video node

The InfiniTalk node's prompt describes the attitude, not the text. What the person says is the wired voice: the text goes in the Voice-over node.

A voice too long for a single pass

Beyond twenty-eight seconds, the node refuses before generating, with no charge. Split into two takes and chain them in the Montage node.

Frequently asked questions

How long can the photo talk?

From two to twenty-eight seconds per generation, about seventy words at a normal pace. The video lasts exactly as long as the wired voice: nothing to set.

Can I use my own voice?

Yes. Upload a recording into a Media node and wire it to the voice port, or clone your voice in the Voice-over node and have it read any text.

Do I need a photo of me, or does a generated character work?

Both. A portrait generated by an Image node, a drawn mascot or an old photo talk as well as a selfie, as long as the face is frontal and the mouth visible.

Am I charged if the generation fails?

No. The price is reserved at launch and returned automatically on a technical failure. The exact amount shows on the button before you click.

What is the difference with the talking avatar?

The talking avatar generates an invented presenter with their voice, in one go, on Veo. Here, YOUR photo talks, with the voice of your choice, and the text changes without regenerating the character.

Can I choose the language?

Yes, in the Voice-over node: the ElevenLabs voices read fourteen languages. The lip sync follows the sounds, whatever the language.

Is the video 16:9 or vertical?

InfiniTalk renders a frame close to the photo, at 480p or 720p. For a precise format, reframe afterwards in the Montage node, which accepts 9:16, 1:1 and 16:9.

Can I make several people talk?

One person per node. For a dialogue, have two photos talk each on its voice, then alternate the shots in a Montage node.

Make a photo talk with AI

Use this workflow

Free account, no card required. The workflow opens pre-filled in your canvas.