August 12, 2026
Making a 10-second AI video ad: the complete case study, prompts and bill included
A real brief: an artisan coffee brand wants a 10-second clip for social media. We build it end to end on Imaginode: key visual, camera settings, Seedance animation, 9:16 variant, with every prompt quoted in full and every credit counted. The honest final bill: about 140 credits, or €1.40.
the brief: Torréfaction Morel, 10 seconds, social media
The client is plausible because it exists everywhere: a small artisan coffee roaster, let's call it Torréfaction Morel, that sells its coffee online and wants a 10-second clip for Instagram and TikTok. No budget for a shoot. The brief fits in three lines: warm workshop atmosphere, a coffee bag front and center, steam, a camera slowly pushing in. Delivery in 16:9 and in 9:16 vertical.
Two years ago, this brief went to a videographer for €800 to €1,500, with a week of turnaround. Today, we're going to handle it in one hour on an Imaginode canvas, for about 140 credits, roughly €1.40 excluding subscription. It's not magic, and we'll see that not everything works on the first try. But that really is the order of magnitude.
One honest clarification before we start: AI video still fails about one generation in three or four. A movement that goes off the rails, one hand too many, steam that looks like fire smoke. And an ugly but technically delivered result stays paid for. So we build the failures into the budget from the start, like a genuine line item.
why a workflow rather than a string of prompts
You could do all of this in a chat interface, chaining prompts and crossing your fingers. The problem is what comes after: the client will ask for a variant, then another, then the Christmas version in November. If your process only exists in a conversation history, every new request sends you back to zero.
On the Imaginode canvas, we're going to build a small machine instead: one node for the idea, one for the art direction, one for the image, one for the camera, one for the video, all linked by cables. Once the machine is built, producing a variant means changing a single node and relaunching. That's the whole difference between cooking a dish and writing down the recipe.
The workflow stays saved automatically, with the last 12 generations kept inside each node. A month from now, when Morel wants the same ad for its new decaf bag, we'll reopen the canvas, swap one image and two words, and the machine will spit out the clip. That month of peace is easily worth today's hour of construction.
step 1: the coffee bag becomes a reference
First building block, and the one beginners skip most often: the Reference node. In it we create "PaquetMorel" with four photos of the real bag from several angles and a short description: brown kraft paper, cream label, black lettering logo, 250 g size. From now on, typing @PaquetMorel in any prompt on the canvas injects the description and attaches the photos automatically. The mention displays in green.
Why this is non-negotiable: without a reference, every generation reinvents a different bag. A label that changes color, a logo that turns into gibberish, fanciful proportions. For an ad, that's disqualifying, the product has to be the real one. The reference doesn't guarantee 100% perfection, but it turns a hopeless fidelity rate into a very workable one.
Second building block, a Style node: art direction reading "artisan workshop, golden hour, light photographic grain, warm brown and cream palette", accompanied by three moodboard images dragged and dropped in, two workshop photos found in the client's image bank and one photo of his actual shop imported through a Media node. These two nodes cost nothing in generation; all they do is feed the ones that follow.
step 2: write the idea in a Text node
Next we drop in a Text node, set to Claude, with this instruction: "write an advertising image prompt in English for a single scene: @PaquetMorel sitting on a wooden workbench in a coffee roasting workshop, steaming cup beside it, coffee beans scattered around, late-afternoon light through a window on the left. One scene only, no embedded text."
Why not write the image prompt yourself? You could. But the Text node produces a denser description than what most of us write off the cuff, and above all it becomes a reusable control point: for the Christmas version, we'll swap "late-afternoon light" for "garlands and winter light" in the instruction, and the whole chain downstream will follow.
A detail that matters: we explicitly ask for "no embedded text". Image models love adding invented slogans in questionable typography, and that's the kind of flaw you can't fix afterward. The legal lines and the tagline will be added later in the editing app, exactly like on a real production.
step 3: the key visual, 1-credit drafts first
The generated text travels down a cable into an Image node. You drop the cable anywhere on the node's body and it snaps onto the right input by itself. The Style node is connected alongside, and its input dot lights up. First money-saving reflex: we draft in Flux Schnell first, 1 credit per image, to validate the composition without burning anything.
Three drafts later, 3 credits spent, the verdict is in: the composition with the window on the left works, but the cup is stealing the show from the bag. We change one line in the Text node, "the cup set back, slightly blurred, the bag sharp in the foreground", and regenerate. That is canvas iteration in a nutshell: you fix the failing part, and everything else stays put.
Composition validated, we move to the serious model: Seedream 5 Lite, 5 credits per image, the champion of references. It's the one that respects @PaquetMorel best. Two generations, 10 credits, and the second one is the keeper: the bag is faithful, the label legible, the light exactly on the axis we asked for. Running total for the key visual at this point: 13 credits, about 13 cents.
step 4: the camera, 85mm, f/1.4, slow dolly-in
A successful key visual isn't enough; you have to decide how the camera will live inside it. That's the job of the Camera node and its five illustrated dials: framing, lens in millimeters, aperture, angle, movement. For Morel, we set: close-up framing, 85mm lens, f/1.4 aperture, slightly high angle, slow dolly-in movement.
Why those exact values? The 85mm is the portrait lens par excellence: it compresses perspective and flatters the subject, in this case the bag. The f/1.4 aperture gives a melted background that isolates the product, precisely the look of high-end coffee ads. And the slow dolly-in creates the feeling of being drawn toward the product, the single most effective cliche in product advertising for the last fifty years.
The node writes the proper English terms at the head of the prompt all by itself, close-up, 85mm lens, f/1.4, and places the movement in a final sentence, the position where video models understand it best. If you've ever typed "cinematic camera" into a prompt hoping for a miracle, these five dials are the grown-up version of that reflex: the exact vocabulary, in the right place, with nothing to memorize.
step 5: the animation, Seedance 2.0 Fast in 720p
Now for the Video node. We plug the key visual into the start image input, one image accepted, that's the rule. We select Seedance 2.0 Fast: about 48 credits for 5 seconds in 720p, and between 60 and 100 credits for 10 seconds depending on settings. Our 10-second clip in 720p reads 76 credits on the Generate button, before any click. No surprise is possible; the price is right there, spelled out.
Why this model and not another among the 47 available? Kling 2.5 Turbo, at around 26 credits for 5 seconds, would have been enough for a simple movement test. Veo 3.1 Fast would have added generated audio, interesting but useless here since Morel will lay down its own music. Seedance 2.0 Fast offers the best product-fidelity-to-price ratio for this kind of slow shot. Its big brother, the full Seedance 2.0, goes up to 1080p and 4K, but 330 credits for 5 seconds in 4K, for a compressed Instagram post, would be throwing money away.
The movement prompt, written in English and polished by the magic wand for 1 credit, fits in four sentences: "Steam rises gently from the cup and drifts across the frame. Dust particles float in the golden light. The coffee bag remains perfectly still and sharp. Slow dolly-in toward the bag, ending in a tight close-up on the label." Keeping the bag motionless is the most important instruction of all: without it, video models love making rigid objects ripple.
first take: failed, second take: nailed
Full transparency about what happened. First generation, 76 credits: the steam is beautiful, the dolly-in is smooth, but at the six-second mark the bag's label starts slowly melting, like a candle. Unusable. And since the clip was technically delivered, it's paid for. Those are the rules of the game, better to know them before rather than after.
We adjust the prompt: "label text stays static and legible during the entire shot" added just before the movement sentence. Second take, another 76 credits: this time everything holds, the steam, the light, the label sharp from start to finish. The 10-second clip is good, and frankly better than what we were hoping for at the brief stage.
One successful generation out of two puts us right at the honest average for this craft. Note that if the failure had been technical, an error on the model provider's side, the refund would have been automatic within seconds. Here, the failure was aesthetic, therefore paid. Real video budget at this point: 152 credits for one approved clip. That's the line item that weighs, and by far.
per-node history, the anti-regret insurance
A useful detour before the variants. Every node keeps its last 12 generations, viewable in place. Our three Flux Schnell drafts, our two Seedream takes, our two videos: all of it is still there, inside the relevant nodes, with no downloads folder to dig through.
It proved useful immediately. Comparing the two Seedream takes, we hesitated about going back to the first one, where the cup had more presence. One click in the Image node's history to recall it, one look, and no, the second one stays the keeper. That kind of back-and-forth, which takes three seconds here, is simply impossible in a chat interface where old images sleep forty messages deep.
The history also changes your relationship to risk. You dare an exotic setting because you know the previous version stays one click away. On client work, that safety net is worth gold: the client who says "actually I preferred the other one" is no longer a catastrophe, it's a click.
the 9:16 variant: one setting changed, not a project redone
Morel also wants a vertical version for stories and TikTok. In a chat, that would mean another complete cycle, with a result probably different from the 16:9. On the canvas, we duplicate the image and video branch, switch the Image node's output format to 9:16, and regenerate the visual: 5 credits with Seedream, which recomposes the scene vertically while keeping the same world, the same light, the same faithful bag thanks to @PaquetMorel.
For the vertical video, we shorten to 5 seconds, the stories format lives perfectly well with short clips, and it halves the cost: 48 credits with Seedance 2.0 Fast in 720p. Same movement prompt, with the dolly-in slightly faster to fit the duration. First generation usable on the first try, this time. That happens too, and nobody complains about it.
Time spent on the variant: nine minutes, stopwatch in hand. Cost: 53 credits. That is exactly the argument for the workflow: the second output costs a fraction of the first, because all the intelligence, references, style, camera, prompts, was already wired in. A third variant, a 1:1 square for the feed, would cost the same the day Morel asks for it.
the final bill, line by line
Let's count everything, without rounding in the convenient direction. Flux Schnell drafts: 3 credits. Seedream 5 Lite key visuals, two takes in 16:9 and one in 9:16: 15 credits. Magic wand on the two video prompts: 2 credits. The 10-second 16:9 video, one paid failure and one approved take: 152 credits. The 5-second 9:16 video: 48 credits.
Grand total: 220 credits, about €2.20, for two clips delivered in two formats. Counting only the main 10-second film, the bill drops to about 140 credits, €1.40, failure included. For comparison, the product photo session alone in the traditional brief would have cost more than one hundred times that amount.
On the subscription side, this project fits very comfortably inside the Starter plan at €13 before tax for 900 monthly credits: it consumes a quarter of them. A creator producing this kind of ad every week will be more at ease on the Creator plan, €42 for 3,100 credits. And for a one-off crunch, top-ups exist, 1,000 credits for €15. The code SIMPLE15 takes 15% off along the way.
what would have blown the budget
First classic mistake: generating directly with an expensive model. Running the drafts in Flux 2 Pro at 7 credits instead of Flux Schnell at 1 credit would have multiplied by seven a line item that must stay negligible. Cheap drafts, then a polished final, is the golden rule, and it counts double for video.
Second mistake: grinding away at 10 seconds. Each 10-second take costs 76 credits; testing the movement on a 5-second clip at 48 credits, or even in Kling 2.5 Turbo at 26 credits, then only going to 10 seconds once the movement was validated, would have saved part of the failed take. We paid for it so you don't have to; keep the lesson.
Third mistake, the worst one: skipping the product reference. Without @PaquetMorel, every generation would have invented a different bag, and we would have burned dozens of credits chasing logo fidelity. The five minutes spent creating the reference are the best investment of the entire session. Without references, a product changes appearance the way a character changes faces: systematically.
the lazy version: ask the assistant for everything
Everything above can be had without building a single node yourself. The bubble in the bottom right of the canvas opens the Imaginode assistant. You write: "build me a workflow for a 10-second product ad: product image from a reference, warm workshop style, 85mm close-up with a dolly-in, Seedance animation in 720p". Cost of the message: 1 credit.
The assistant builds the complete machine: nodes chosen, prompts written in English, connections wired, all of it added to the canvas in one click and laid out cleanly in columns, with no overlap. All that's left is creating the reference with your photos, proofreading the prompts, and clicking Generate. The hour of construction described in this article compresses into five minutes.
Our advice, having done it both ways: build your first workflow by hand, following a case like this one, to understand what each cable carries. Then switch to the assistant for all the ones after. Since it sees your open canvas, you can also ask it along the way why a result disappoints or which model to pick: it advises on your actual case, not in theory.
publish, measure, repeat
The two clips go off to Morel, who adds his music and his tagline in his usual editing app, then publishes. The workflow, meanwhile, stays on the canvas, saved automatically, ready for the next request. A true anecdote from the daily life of a creator: the "new decaf bag" version was requested ten days later, and delivered in twelve minutes from a phone, on the train, the generation running with the screen locked.
Because yes, all of this works from a phone's browser, full touch canvas included, and picks up identically on a computer. For relaunching a node or approving a take, it's perfect. For building a big workflow from scratch, let's stay honest: a mouse and a large screen remain more comfortable. Mobile is an excellent cockpit, not the best construction workshop.
The logical next step if this case study got you interested: the blog article on the Camera node, to learn how to choose your focal lengths and movements like a cinematographer, and the one on workflows to understand the mechanics of the cables in depth. Imaginode's image-to-video template is waiting for you in the meantime: load it, replace the prompt, and your first ad is under 100 credits away.

Balance