The full pipeline of an explainer channel: a subject in, a script written by a language model, a voice over, three shots and a ten second edit. Visible here without an account.
The whole workflow is here: the text node writes the script, the voice over reads it, the art direction holds the three shots, the video node renders the edit. Pan, zoom, open the nodes.
At a glance
What the workflow produces
3 images in 16:9 · 1 video of 10s in 720p, 16:9 · 1 voice-over
Canvas structure
10 nodes wired by 7 connections
Models wired in
Kimi K2.5, OpenAI TTS HD, Seedream 5 Pro, Seedance 2.0
Cost of one full run
about 139 credits, roughly €1.39
A faceless channel lives on one thing: producing fast, often, without ever filming yourself. A single prompt box is not enough, because this kind of video is not an animated image but a production chain: a subject, a text, a voice, shots, an edit. This workflow shows the five links side by side, and that is exactly what you get when you open the template.
The script node is the only one in the product to carry a language model: you type a subject, it returns the voice over text, and that same text runs into the voice node. Nothing to copy between two tools, and nothing to redo when the subject changes. The voice over stays in its own node rather than inside the video: you can replay it as often as you like without paying for the edit again.
The first text node holds one line: the subject of the video. It is the only place to change when you produce the next episode.
2
Let the script write itself
The second text node carries a language model and a writing instruction. Count about thirty words for ten seconds: beyond that, the voice runs after the images.
3
Lock the art direction
The art direction node describes the material, the palette and the light once, and feeds all three shots. Without it, the three images do not look like they come from the same video.
4
Edit in a single render
The video prompt describes four numbered shots, hard cuts, two and a half seconds each. The edit comes out silent: the voice over sits on top and stays editable on its own.
Variations and use cases
The same canvas, tuned for different needs: each variation is two or three lines of prompt away.
Mystery and true story channel
The densest faceless niche: a cold case, a disappearance, a puzzle. The script Text node writes the narration from your subject, the voice-over reads it, and the three shots show places rather than people. Faces are what give a generated video away fastest.
Science explainer channel
Same chain, different tone: a question up front, an explanation in three beats, a conclusion. Always read what the language model wrote before generating the voice: on a technical subject it states false things in a perfectly even tone.
Vertical Shorts from the same script
The script lives in its own node: duplicate the Video node in 9:16, keep three shots instead of five, and you have the short format without rewriting a line. That is the main gain of a canvas over a one-field generator.
Faceless video in another language
Write the subject in the language you want: the language model drafts in it and OpenAI TTS HD reads it properly. The same canvas then produces one version per market, with identical shots and a different voice-over.
A series on a regular schedule
A faceless channel lives on consistency. Keep one canvas per series: only the Subject node changes between episodes, the art direction and the voice stay, and the visual identity builds itself.
Common mistakes
Generating the voice before proofreading the script
A voice-over rendered on an unread script has to be redone in full for one word. Read first, generate second: the voice lives in its own node precisely so it can be redone alone.
Three shots without a shared art direction
Without the Style node wired into all three images, the video looks like three different videos glued together. That is the flaw that loses viewers after four seconds.
An eighty-word script for ten seconds
Count thirty words for ten seconds. Too long a script gives a rushed voice, and the edit trails the narration from start to finish.
Why is the voice over not generated inside the video?
Because a voice baked into the render means paying for the whole video again to change one word. In its own node it costs a few cents to redo, and it is what dictates the length of the shots.
What does a video like this cost?
Three images, a voice over and a ten second edit come to around one and a half euros of credits. The model by model breakdown is on the calculator.
Can it run longer than ten seconds?
Yes, by chaining several video nodes on the same canvas, each starting from the last frame of the previous one. The template shows one to stay readable, the mechanism is the same beyond that.
Can a faceless channel be monetised on YouTube?
Yes, provided you add something of your own: an original script, an edit, a voice, a point of view. YouTube penalises repetitive, inauthentic content, not the absence of a face or the use of generation tools.
Which language model writes the script?
The template's Text node runs Kimi K2.5 at one credit per generation. Any other language model in the catalogue can replace it in the same node.
Can I produce one video a day?
Yes: once the canvas is built, an episode means changing the subject and hitting generate, so roughly ten minutes and about 1.40 euros of credits for ten edited seconds.
Do I have to disclose that the video is AI generated?
YouTube asks you to flag realistic synthetic content that could mislead, through the checkbox in the upload flow. A documentary illustrated with generated images gets flagged, and that affects neither monetisation nor reach.