The « I'm inside the game » reel
The format that inspired this workflow: a move filmed vertically, a game world around it, five seconds. The node renders straight to 9:16 with the world's sound, the result posts as is.
A clip of you shot on a phone enters the canvas, and Wan 3.0 renders the same person, the same move, inside a video game level. Temple ruins, cube world, neon rooftop with the HUD: the whole workflow is visible without an account, with all three renders.
Real workflow, real results. Pan and zoom freely.
Use this workflowThe real workflow: an ordinary clip shot in a hallway, three Video nodes on Wan 3.0 that receive that clip as a reference through the video port, and three prompts that describe the person, then the world they end up in.
At a glance
The idea comes from those videos where someone ends up « for real » inside a video game: the person is filmed, the world around them is the game's, and that contrast is what makes people laugh or stare. Here there is no green screen and no compositing software: a clip shot on a phone, in a hallway, enters through a node's video port and the same person comes out jumping from one stone platform to the next above a chasm.
The clip is given as a reference, not edited. Wan 3.0 reads it the way it reads a character sheet: the face, the bun, the yellow hoodie, the jump, the grin at the end, and it renders a new video where all of that is kept inside the world the prompt describes. So the prompt has two halves: the person first, stating that they remain a real filmed person, then the world and the light it casts on them.
Three nodes receive the same clip, because a world is chosen by comparing: the jungle temple ruins of an adventure game, a cube world where everything is made of blocks except her, and a neon rooftop from an action game with the health bar and the minimap drawn over the picture. Each world costs sixty credits for five seconds in 720p with sound, all three together one hundred and eighty, a little over two dollars.
Film a move that means something in a game
Jumping, balancing, aiming, spinning around. Five to ten seconds, wide framing, the whole person in the frame. A clip where nothing happens gives a world where nothing happens.
Replace the demo clip with yours
The Media node accepts a video imported from the phone. The three Video nodes are already wired to it: nothing to reconnect. Duration, format and resolution are set on the node, the reference does not impose them.
Describe the person, then the world
Say they are the one from the reference video, their face, their clothes, their move, and that they stay a real filmed person. Then the ground they now walk on, the world's light on them, the interface if you want one.
Generate, compare, keep the world that holds
Three renders side by side, one will be better. Rerun only that one with an adjusted prompt: each node is paid separately, the other two stay put.
The same canvas, tuned for different needs: each variation is two or three lines of prompt away.
The format that inspired this workflow: a move filmed vertically, a game world around it, five seconds. The node renders straight to 9:16 with the world's sound, the result posts as is.
Everyone films five seconds, everyone goes through the same world, and the edit puts them side by side. The Edit node assembles the renders in the browser, without credits.
The world you describe is yours: your colours, your interface, your enemies. A real person jumping inside it says more about scale and mood than a screenshot.
One source, ten Video nodes, ten prompts. The clip is never refilmed, the world changes at every node, and the comparison happens on the canvas.
Describe the world without naming the game: the cubes, the ruins, the stone blocks. The child stays themselves on screen, which matters when the video is shown to the grandparents.
Wan 3.0 already outputs the sounds of the world, but game music on top is generated with the make a song workflow and laid into the Edit node.
Standing still, arms at their sides, the render is a photo in a set. The move makes the game: jumping, balancing, aiming. Five seconds of movement beat ten of posing.
Without that sentence, the model reads the video as a mood and invents someone else. The prompt starts with the person: the one from the reference video, their face, their clothes, their move, and they stay a real filmed person.
A game title gives the model little and raises a rights problem. Describe what is seen: the ground, the light, the objects, the interface. « Textured cubes, a square sun » is enough.
A torch-lit temple without warm light on the hoodie stays a pasted background. Always ask for the set's light on the clothes and the face: that is what makes them part of it.
Yes, that is the heart of the effect, and the demo shows it across three worlds: the face, the bun, the yellow hoodie and the jump are identical in all three, down to the grin at the end. The move is kept because the whole video serves as the reference, not a single frame.
Yes, on Wan 3.0, which has never refused a clip, and it was measured. Seedance 2.5 does better when it accepts, since it edits your clip itself instead of redrawing it, but its filter sometimes refuses a person, real or generated: it refused a test clip for this demo. The node warns you before generating when a model refuses real people.
Yes, by describing it as something drawn over the picture: the health bar, its colour, its corner, the minimap, the score. The third render of the demo does exactly that. Do not name a game, describe what is seen.
Sixty credits for five seconds in 720p with sound on Wan 3.0, that is seventy-two cents. The three worlds of the demo add up to one hundred and eighty credits. The price shows on each node before you click.
One to fifteen seconds as a reference. The output has the duration set on the node, two to thirty seconds: a five-second clip rendered as five seconds keeps the whole move, that is the demo's setting.
No. The format is chosen on the node, independently of the reference video: a clip shot in landscape can come out vertical, and the other way round. The demo is 9:16 because that is the reels format.
Yes, the node lists every model that accepts a video input. Seedance 2.5 edits your clip itself, the most faithful option when its filter accepts it. Wan 3.0 Prime renders finer for forty percent more. Grok Imagine Video accepts the clip too but redraws the person: other clothes, another framing, it is no longer you.
Yes, and it is the natural case: you film, you import, you run. The create on mobile guide describes the canvas gestures.
Put yourself in a video game with AI
Use this workflowFree account, no card required. The workflow opens pre-filled in your canvas.