The canvas and nodes
Text, image, video, media, reference, style: how to wire nodes into workflows.
Node types
Text holds an idea or prompt, and can even write it for you with a language model (GPT-5, Claude, Gemini…). Image and Video generate content with AI. Media imports your own files. Camera sets how the scene is framed. Reference remembers a recurring character or object. Style enforces an art direction.
Every node moves freely; delete it with the cross or the Delete key. Undo and redo with Ctrl+Z / Ctrl+Y.
Wiring nodes
Drag a link from a node's right handle to another node's left handle. A Text node wired to an Image's “prompt” port adds its text to the prompt: a blue inset in the node confirms what is taken into account. An Image wired to a Video becomes its first frame. An image Media serves as a visual reference.
Some video models also accept an end image: the “end image” port appears automatically when the model supports it.
@References and styles
Create a Reference node with a name, a description and photos: your character. Then type @ in any prompt: your references pop up, and the chosen mention is shown in neon green. At generation time the description is injected into the prompt and the images are attached (with an edit-capable model such as Seedream 5 Lite or GPT Image Mini).
The Style node works the same but through wiring: connect it to your Image and Video nodes to keep the same art direction everywhere.
Combining
A complete workflow example: Text → Image (Flux 2 Klein) → Video (Kling), with one Camera node wired to both for a consistent look and a shared Style. Each Image and Video node keeps its generation history: click a thumbnail to go back. Everything saves automatically.
The Retouch node
The Retouch node is a small layer workshop on the canvas. Wire Image or Media nodes to its Layers port: each image becomes a layer you move, resize, rotate, fade and blend with the others. The eraser makes an area transparent (to cut out a subject or open a window onto the layer below), the Restore brush brings it back. Add text on top, choose the document size and background, then export: the PNG keeps its transparency, everything is computed in your browser, with no credits.
Every image layer has two AI actions: retouch an area (highlight, describe, Nano Banana replaces) and upscale (Topaz, ×2 or ×4), at the same price as on the other nodes. One click returns to the previous image. The node's output wires anywhere an image does: a video's start frame, a reference, a montage.
The video tools: motion, extension, lips, background
The Video node does more than generate: depending on the model you pick, it transforms an existing video. Kling 3.0 Motion Control takes the gesture of a video (the “Motion video” port) and has the character of an image (the image port) replay it: a dance, a walk, a choreography, on any face. Grok Imagine (extend) generates the continuation of a video from its last frame, in its format: you set how many seconds to add and describe what happens next. The continuation comes out on its own, to chain in the Montage node.
Sync Lipsync 3 re-times the lips of a video to a wired voice (the voice port): another language, another script, on the same shot. InfiniTalk does the same from a photo: the face talks for the whole duration of the voice, expressions included. Bria video background removes the set from a video and keeps the subject and its sound: a transparent webm for your edits, or an mp4 on a green, black or white background.
These five tools are billed on the source measured server-side, never on a declared duration: the node reads the length of the wired video or voice and shows the exact price before you click. A model only shows the settings it truly honors: when a selector is missing, the model does not take it into account.
Reading a PDF in a Text node
Drop a PDF on the canvas: a Media node receives it, like an image or a video. Wire it into the “Image or PDF to analyze” port of a Text node set to a model that reads documents (GPT, Claude, Gemini, Kimi K3, Grok 4.6, DeepSeek V4 and the others marked as such): the model receives the whole document and answers your instruction, summary, extraction, rewrite, content plan, illustration prompts.
The price follows the size of the document: the node reserves the cost when the file is submitted, then settles on the number of tokens actually read, and the difference returns to your balance. Up to three documents per node. The answer feeds the Image and Video nodes wired behind it directly: a report becomes a carousel, a manual becomes a video tutorial.
My characters
A well-filled Reference node (name, description, photos) is worth keeping. The “Save to My characters” button stores it in your account. In any other project, “My characters” opens the list and fills a Reference node in one click, name, description and images included. Up to thirty characters per account, eight images each.
An edited character updates from the node that loaded it, and deletes from the list. That is what lets you keep the same mascot, the same presenter or the same product from one week to the next, without rebuilding the card.
The 3D Object node
The 3D Object node turns an image into a 3D file. Wire an image (an uploaded photo or the render of an Image node): Tripo 2.5 or Trellis 2 make a textured GLB of it, the format read by Blender, Unity, Unreal and web viewers. The model displays in the node, rotates with a finger or the mouse, and downloads.
Two renders are then made in your browser, at no cost: a five-second turntable (video port) and a still view (image port). The turntable wires as a reference into a Video node, Wan 3.0 or Seedance 2.5, which keeps the shape of the object while staging it. The view wires anywhere an image goes in. A whole object on a plain background gives the best volume; a scene or a subject cut by the frame gives a worse one.