August 12, 2026
Directing AI like a cinematographer: lenses, framings and movements that change everything
Image and video models learned from millions of photos captioned with cinema vocabulary. Talk to them in 85mm, f/1.4 and low angle, and they obey. Imaginode's Camera node writes those terms for you, at the top of the prompt, right where they carry the most weight.
the word your prompts are missing
Take two prompts. The first: "portrait of a red-haired woman in a coffee shop". The second: "close-up, 85mm lens, f/1.4, portrait of a red-haired woman in a coffee shop". The first produces a decent image, framed at random, with a background fighting the subject for attention. The second produces a tight portrait, background melted into patches of light, sharp gaze. Same subject, eight words apart, two images that have nothing to do with each other.
Those eight words are not magic spells. They are the exact terms a cinematographer would give a camera assistant before rolling the shot. And it turns out AI models understand them remarkably well, for a precise reason we're about to break down.
This article walks through every setting with its concrete effect on the image, shows how Imaginode's Camera node writes all of these terms for you, and ends with the real difficulty of multi-shot work: keeping the same direction from one shot to the next.
why models understand cinema vocabulary
Image and video models were trained on enormous quantities of captioned photos and clips. Now, who captions their images carefully? Photographers, stock image libraries, gear review sites. Their descriptions are packed with technical mentions: "shot on 85mm", "f/1.4", "low angle shot", "dolly in".
As a result, those terms are anything but decorative for the model: they are statistically tied to very distinctive looks. Writing "85mm" summons every 85mm portrait the model has ever seen, with their compressed perspective and soft backgrounds. That's infinitely more reliable than "beautiful professional photo", which corresponds to no particular look at all.
The practical consequence is thrilling if you come from photography, and frustrating for everyone else: an effective prompt reads like a shooting call sheet, in English. The good news is that you don't need to learn any of it by heart.
the Camera node: five dials, zero jargon to memorize
On the Imaginode canvas, the Camera node replaces memorization with five illustrated dials: framing, lens in millimeters, aperture, angle and movement. You turn the dials while looking at the thumbnails, and the node writes the matching English terms for you. No need to know that the French "plan américain" is called a "cowboy shot".
One detail that changes everything: those terms are inserted at the top of the prompt, before your scene description. Models weight the beginning of a prompt more heavily, so a setting placed in first position is respected far better than the same setting drowned at the end of a sentence. The movement, for its part, is added as an explicit closing sentence, the form video models follow best.
Then plug the Camera node into an Image or Video node like any other input: one cable, the port lights up, and it's in place. A single Camera node can feed several generators at once. We'll come back to that, because it's the secret weapon of multi-shot consistency.
the six framings and what they say
Framing is the first decision, the one that defines what the viewer looks at. The dial offers six. The extreme close-up isolates a detail, a pair of eyes, a hand on a door handle, and manufactures tension. The close-up frames the face: it's the framing of emotion, the one that makes a character exist.
The medium shot cuts at the waist and suits dialogue and gestures. The cowboy shot, cut at mid-thigh, comes from westerns where you needed to see the revolver, and remains the natural framing for someone in the middle of doing something. The wide shot places the character in their setting. The establishing shot crushes the character inside a huge environment: perfect for opening a story.
The beginner mistake is on display every single day: generating everything in medium shot, the default framing models fall back on when you specify nothing. An entire sequence shot at the same distance from the subject is flat, whatever the model's talent. Alternate: an establishing shot to set the place, a cowboy shot for the action, a close-up for the reaction.
the lens: 14mm distorts, 135mm compresses
The lens dial runs from 14 to 135mm, and that number changes the very geometry of the image. A 14mm takes in a huge field and stretches perspective: lines rush away, foregrounds become monumental. Gorgeous for architecture or a cramped interior, cruel to a face, which it warps like a funhouse mirror.
The 24mm and 35mm are the reportage lenses: wide enough to give context, tame enough not to distort. The 50mm roughly reproduces the perspective of the human eye, hence its nickname of normal lens. It's the neutral choice, the one that tells you nothing beyond the scene itself.
Beyond that, telephoto compresses. A 135mm visually pulls the planes of the image together: the mountain seems to sit right behind the character, the line of cars becomes a solid wall. That compression, combined with background blur, produces the most "cinema" images of the lot. Try the same street scene at 24mm and then at 135mm: you won't believe it's the same place.
the 85mm, best friend of portraits
If you remember one number from this article, make it 85. The 85mm is the portrait focal length par excellence: long enough to flatten features slightly, which flatters practically every face, short enough to keep some context behind the subject. Noses look thinner than at 24mm, proportions look more true to life.
AI models have seen millions of portraits captioned "85mm", so the term is one of the most powerful in the whole vocabulary. On an Image node, the combination of close-up, 85mm and a wide aperture almost guarantees a portrait that's sharp on the eyes and melted behind, the look people associate with high-end wedding photography.
True story from a product page project: the same portrait prompt, in Flux Pro 1.1 at 6 credits, first with no settings and then with the Camera node set to 85mm f/1.8. The first version was usable. The second was mistaken for a studio photo by the client. Twelve credits total for the comparison, and a lesson that stuck.
aperture: f/1.4 isolates, f/16 shows everything
The aperture dial goes from f/1.4 to f/16, and it drives one visible thing: depth of field. Simple rule, wide aperture equals blurred background. At f/1.4, only the subject is sharp, the background dissolves into soft washes and circles of light. At f/16, everything is sharp from the foreground all the way to the horizon.
f/1.4 is the isolation tool: a face detached from a cluttered background, a product standing out on a crowded table. It's also an admitted cover-up: when the model generates an approximate background, the blur hides it elegantly. Plenty of creators use it for exactly that without ever saying so.
Conversely, f/8 or f/16 serve wide scenery and architecture, where every plane has to stay readable. Watch out for coherence: asking for a wide mountain shot at f/1.4 is contradictory, and the model then answers however it pleases. Wide aperture on tight shots, closed aperture on wide shots: that one reflex settles 90% of cases.
angles: making a subject heroic or crushing it
Eye level is the neutral angle, the documentary angle. The moment the camera leaves that height, it takes sides. The low angle, camera down low looking up, enlarges the subject and makes it heroic: it's the angle of superhero posters and CEO photos on magazine covers. A few inches of viewpoint, and your character dominates the scene.
The high angle does the opposite: it crushes, isolates, weakens. A character filmed from above in a big empty room says loneliness without a single word. The overhead angle, shot straight down, turns the scene into a graphic composition, heavily used for food and neatly arranged desk shots.
That leaves the Dutch angle, the deliberately tilted horizon, which installs unease and instability: chases, crisis scenes, moments where something is off. Use it rarely, precisely because it gets noticed. An entire teaser shot in Dutch angle mostly delivers seasickness.
why all of this belongs at the top of the prompt
A prompt is not read like a contract, with every word carrying equal weight. Models give more importance to the beginning of the text, and that weighting fades toward the end. A "low angle shot" in first position will be respected; the same term at word forty will sometimes be ignored, especially if the scene description is rich.
That's exactly the anatomy the Camera node enforces: technical terms first, framing, focal length, aperture, angle, then your scene description, your character @mentions in green, your style. And the camera movement as an explicit final sentence, because video models handle a movement instruction better when it's phrased as a complete action rather than an adjective.
If you write your prompts by hand, copy that structure. And if your prompt is a rough draft in your own words, the magic wand on the Image and Video nodes rewrites it into structured English via the Kimi model for 1 credit, preserving the @mentions. It respects this same hierarchy: technique up top, scene next, movement to close.
before and after: the same scene, directed and not
Test scene: "a baker pulls a batch of bread from the oven, morning light". Without direction, models typically deliver a medium shot at eye level, everything sharp, decent lighting, stock-photo composition. Technically clean, instantly forgettable. It's the statistical average of every bakery seen during training.
Pass number one with the Camera node: close-up, 85mm, f/1.4, eye level. The result tightens on the focused face, the steam of the bakehouse melts into a halo, the bread becomes a golden smudge in the foreground. Pass number two: wide shot, 24mm, f/8, slight low angle. This time the whole shop exists, the shelves rush away in perspective, the baker becomes the pivot of a place.
Neither version is "better". They are two different intentions, and that's exactly the point: without settings you had no intention, you had an average. Cost of the full experiment in Imagen 4 Fast: 9 credits for three images, about €0.09. Camera direction is the cheapest lever in all of AI creation.
the video movements, one by one
In video, the fifth dial comes into play. The static camera is the default choice to own proudly: it lets the scene live and removes one source of error for the model, since a locked shot fails less often than a shot with a complex move. The dolly in closes on the subject and creates immersion or threat; the dolly out reveals context or signals the end of a scene.
The pan sweeps the scene on a pivot, ideal for discovering a set or following a movement across the frame. The crane up rises above the scene, the movement of film endings. The orbit circles the subject and turns it into a precious object: heavily used in product ads, spectacular but demanding, it's the movement models botch most willingly.
The handheld camera adds a human shake, precious for action scenes or mock documentary. The slow zoom, finally, tightens the frame imperceptibly and installs a quiet tension: the most discreet movement and probably the most reliable one in generation. Plug the Camera node into the Video node, and the movement sentence is appended at the end of the prompt.
which movement for which intention
The right movement depends on what the shot should make the viewer feel, not on what looks impressive. Opening a teaser? Crane up or dolly out, which establish the scale of the world. Presenting a product? Slow orbit or slow zoom, which install precision. Following a worried character? Handheld camera pushing in, and the discomfort is immediate.
A sobriety rule, learned the expensive way in credits: one single movement per 5-second shot. Prompts that ask for a rising orbit with a zoom produce drunken trajectories. Current video models execute a simple, clearly named movement well, and get lost the moment you load them up with choreography.
And remember what the static camera can do. On a beautiful shot with wind in the leaves and a character breathing, the absence of camera movement is a directing decision, not laziness. Roughly a third of the shots in a professional film are static. Your AI videos can follow the same proportion.
one Camera node for all your shots
The problem shows up as early as the second shot: your shot A is at 35mm at eye level, and without thinking you generate shot B at 85mm in low angle. Cut together, it smells like a collage, without anyone being able to say exactly why. Real films hold together because a cinematographer imposes the same optical choices on an entire sequence.
On the canvas, the solution is structural: a single Camera node can feed several generators. Create one node set to your visual identity, say 35mm, f/2.8, eye level, and plug it into your three Image nodes and your three Video nodes. Six shots, one single source of optical truth.
The day you want to change direction, you turn one dial on one node and regenerate. The history of each node's last 12 generations lets you compare the old version and the new one before deciding. Add a Style node for the palette and a Reference node for the character, and you're holding a pipeline that looks like a real film crew.
camera, style, reference: the complete film crew
The Camera node handles the optics, but a coherent film stands on three pillars. The second is the Style node: a written art direction plus moodboard images, plugged into your generators to impose the same palette and the same texture on every shot. The third is the Reference node, which keeps your characters recognizable from one shot to the next.
The Reference node deserves two sentences of instructions. You give it a name, a description and photos; after that, a simple @mention in any prompt, displayed in green, injects the description and attaches the photos automatically. Without it, your heroine's face changes at every generation, and no dial on the Camera node can do a thing about it.
The typical assembly of a sequence therefore takes very few nodes: one Camera plugged in everywhere, one Style plugged in everywhere, one Reference per character, then one Image or Video node per shot. If wiring all that intimidates you, the assistant in the bottom right builds the complete workflow and adds it to the canvas in one click, for 1 credit per message.
what direction can't save, and what comes next
Let's be frank about the limits. The Camera node makes the look more reliable, it does not fix the models' weaknesses: hands remain dicey, physics remains approximate, and in video one generation out of three or four still goes straight to the trash, perfect settings or not. A delivered but failed shot is still billed; only technical failures are refunded automatically.
What camera direction mostly eliminates is the bad half of the failures: the ones that came from a framing you endured rather than chose. Combined with the method of validating every shot as an image before animating it, detailed in our article on your first AI video, it turns a lottery budget into a production budget.
Concrete next step: open a canvas, drop in an Image node at 1 credit on Flux Schnell, a Camera node, and generate the same scene under three different framings. Three credits, ten minutes, and you'll know by instinct what this article just explained to you. The video course academy built into the documentation shows the exercise in real conditions if you'd rather watch before trying.
