September 5, 2026
AI video longer than 10 seconds: the real durations model by model and the method that costs half as much
How far each video model really goes, why a 30 second generation costs six times a 5 second one, and how to assemble a one minute film by stringing short shots together with free editing.

Frank HoubreFounder of Imaginode
The limit everybody discovers on their first video
Most video models stop between 5 and 10 seconds, a few climb to 15, and only three go all the way to 30 seconds in a single generation.
You have just generated your first shot. It looks good, the movement works, the light is there. And it lasts five seconds. You look for the slider to take it up to a minute, and it does not exist.
It is the first disappointment for everybody arriving at AI video, and it is legitimate: we are sold films and delivered shots. Except the constraint is not the one people think, and the solution has nothing technical about it. It is a hundred years old and it is called editing.
We are going to look at this together, with the exact durations of every model in the catalogue, the real prices in credits, and above all the comparison that hurts: what a minute of video costs in a single generation versus the same minute cut into shots. The gap will surprise you.
The real durations, model by model
Wan 3.0, Wan 3.0 Prime and Seedance 2.5 climb to 30 seconds, Flux 3 Video to 20, Kling V3, Seedance 2.0, Wan 2.7 and Grok Imagine to 15, while Veo 3.1 caps out at 8 seconds.
Here is the state of the catalogue, because those ceilings are written nowhere else in readable form. The longest first: Wan 3.0 and Wan 3.0 Prime accept 2 to 30 seconds, Seedance 2.5 4 to 30. Those are the only three that reach half a minute.
Just below, Flux 3 Video goes from 5 to 20 seconds. Then the pack at 15: Kling V3 from 3 to 15, Seedance 2.0 and Seedance 2.0 Fast from 4 to 15, Wan 2.7 from 2 to 15, Grok Imagine Video from 1 to 15.
And the short ones, which are often the best: Kling 2.5 Turbo offers 5 or 10 seconds, with no value in between. Veo 3.1 Lite and Veo 3.1 Fast give you the choice between 4, 6 and 8 seconds, not one more. The full catalogue is on the models page, with the price moving live according to your settings.
Why the models stop so short
Compute cost rises with duration and coherence degrades: beyond a few seconds, a model loses track of its own scene.
Two reasons, one economic and one technical. The economic one is simple: generating video means producing twenty four images per second that all have to hold together. The compute climbs fast, and providers bill by the second because that is what it costs them.
The technical one is more interesting. A video model holds its coherence over a limited window. Past a certain point, faces slowly transform, clothes change colour, an object left on the table disappears. You have surely already seen a twenty second AI video where something drifts slowly without your being able to say what.
Bear in mind that the 5 or 10 second ceiling is not a commercial slight. It is often the zone where the model is genuinely good. Beyond it, you pay more for a result that holds together less well, which is the worst of both worlds.
The price of a long generation: the nasty surprise
Billing is proportional to the second: thirty seconds cost exactly six times five seconds, with no volume discount whatsoever.
Here is the point people discover after the fact, and it is the real subject of this article. On most models, the rate is pro rata to the duration. There is no sliding scale, no package, no wholesale price.
The exact figures, on Wan 3.0. Five seconds in 720p cost 60 credits. Ten seconds in 1080p, 240 credits. And thirty seconds in 1080p, 720 credits, about 7 euros for half a minute. In 720p that same half minute drops to 360 credits, in 480p to 180.
On Seedance 2.5, thirty seconds in 720p also cost 720 credits, against 240 for ten seconds. Flux 3 Video asks 696 credits for its 20 seconds in 1080p. Kling V3 is at 303 credits for ten seconds in 1080p. At that rate, one failed generation on a thirty second shot and you have just thrown seven euros out of the window.
The comparison that settles it: thirty seconds, two methods
Six five second shots on Kling 2.5 Turbo cost 156 credits against 360 for a single thirty second generation in 720p: less than half, for a more watchable result.
Take a thirty second film and do the sums both ways. Method one, the single generation: Wan 3.0, thirty seconds, 720p, 360 credits. One click, one wait, and one thirty second sequence shot.
Method two, six five second shots on Kling 2.5 Turbo in 720p, at 26 credits each. Total: 156 credits. With Seedance 2.0 Fast at 48 credits for five seconds, you climb to 288 credits, still below. And the editing that assembles it all costs nothing, we will get to that.
But price is not even the main argument. With six shots, you miss one shot and you rerun that one for 26 credits. With a single generation, you miss the twenty second second and you pay for all thirty again. That asymmetry is what counts when you really work, not the initial saving.
What cinema has been doing for a hundred years
The average shot length in a recent film is a few seconds: nobody watches thirty seconds of a fixed shot, and AI simply forces you to respect that grammar.
Look at any advert, any trailer, any music video. Count the shots. In thirty seconds of advertising you will often find a dozen. The long take exists, it is magnificent, it is rare, and it is reserved for precise moments.
A memorable video is not a beautiful image that lasts, it is an emotion that evolves. And an emotion evolves through the cut: you show the face, then what it is looking at, then the trembling hand. Three shots of three seconds tell infinitely more than nine seconds fixed on the same frame.
So yes, the constraint of the models is imposed on you at first. But it pushes you exactly where you needed to go. Those producing the best AI videos today are not chasing the thirty second shot, they are stringing together shots of five.
The Editing node, and the fact that it costs nothing
Editing assembles the clips, trims them, adds cuts or dissolves and a sound track, and the export is computed on your device without spending a single credit.
This is the node that changes the practical stakes of this whole article. You plug your Video nodes into its Clips port, or you pick straight from the project media, and you get a timeline. Each clip is trimmed with a start and an end. Between two clips you choose a straight cut, a dissolve or a fade to black.
It also accepts still images, with an adjustable duration, which lets you slip a title card or a product photo between two animated shots. And it carries separate audio tracks, with an offset and a volume per track: your voice over on one side, the music on the other.
The important point: the export is free. The computation happens in your browser, on your machine, and no credit is debited. So you can re edit your film fifteen times, test three shot orders, change a transition, without it costing you anything. Just keep the tab open during the export.
The complete method, from shot list to export
Cut into shots first, validate each shot as a still for a few credits, only generate video from validated images, then assemble.
Step one, on paper or in a Text node: cut your thirty seconds into shots. Six shots of five, or eight shots of three to five. Write down what each one shows. That is what is called a shot list, and it takes ten minutes.
Step two, and this is the one that saves the most: validate each shot as a still image before paying for video. A Flux 2 Klein image costs 3 credits, a video shot costs 26 to 96. You frame it, you adjust the light, you validate the character, all at image prices. I detailed this method in the AI storyboard before shooting.
Step three: each validated image becomes the starting image of a Video node. The model animates what you have already approved instead of inventing a frame. Step four: it all goes into the Editing node. By that stage, the bulk of the creative work is already done.
Keeping continuity from one shot to the next
A Reference node for the character and a Style node plugged into every shot are enough for the six clips to belong to the same film.
The real risk of editing by shots is that it does not look like a film but like six different videos glued together. The character's face changes, the light jumps, the grain is not the same. That is where the canvas becomes useful rather than pretty.
A Reference node holds the character: you give it three photos and a name, then you call it with an @mention in each of your six prompts. The subject is covered in detail in consistent character in AI, and it is the most important brick if a human appears in more than one shot.
A Style node plugged into the six nodes holds the art direction: same palette, same material, same treatment of light everywhere. And a shared Camera node keeps a coherent lens and aperture. Three nodes for six shots, wired once.
The end frame: chaining two shots with no visible cut
Several models accept an end image in addition to the start image: the next shot can therefore begin exactly where the previous one stops.
There is a trick for the moments when you really want continuity rather than a cut. Some models handle an end image in addition to the start image: you give them the beginning and the arrival, they manufacture the path between the two.
In the catalogue, Seedance 2.0 and Seedance 2.0 Fast, Seedance 2.5, Kling V3, Wan 3.0 and Wan 3.0 Prime, MiniMax H3 and Veo 3.1 Fast accept an end image. The port appears on its own on the node when the selected model can do it, no need to hunt for a hidden setting.
The technique that follows: take the last frame of your shot one and make it the starting image of shot two. You get two five second clips that follow on without a break, for the price of two short clips. That is how you build a twenty second movement that holds up.
Sound: one track, not six
Generating audio on each shot produces six ambiences that contradict each other in the edit: better to have silent shots and one sound track laid over the whole thing.
A classic trap, and it costs you twice. Many models offer an audio toggle, and it is often on by default because a silent video is noticed immediately. On an isolated shot, that is fine.
On six shots destined to be edited together, it is a bad idea. Each generation invents its own ambience, and in the edit you hear six different rooms cutting into one another. And audio is billed: on Veo 3.1 Lite, eight seconds in 720p cost 29 credits without sound and 48 with. Six times, that is 114 credits of difference for an unusable result.
Generate silent, then lay the sound over the complete edit. A voice over, some music, a few sound effects placed at the right points on the Editing node's audio tracks, with their offset and their volume. The film gains a unity that separately scored shots will never have.
When a long generation is justified anyway
The sequence shot keeps its interest on a continuous movement with no possible cut, a tracking shot or a gradual transformation, and there you have to accept the price.
I am not going to tell you long generation is never useful, that would be dishonest. There are cases where it is the only answer. A continuous tracking shot through a set, a transformation that has to be seen without interruption, a dance, a shot where cutting would destroy the very effect.
In those cases, take Wan 3.0 or Seedance 2.5 and accept the rate. But do it knowingly: validate the framing as an image first, drop the resolution during your tests, and only go up to 1080p on the version you know is good. Thirty seconds in 480p cost 180 credits on Wan 3.0 against 720 in 1080p, and to validate a movement 480p is plenty.
Incidentally, that is also the logic of the rest of the catalogue: explore at the floor rate, finalise at the high one. The detail of this arithmetic is in what an AI image or video really costs, and the public calculator lets you simulate your own film before spending a credit.
What to remember
Thirty seconds in short shots cost less than half a single generation, are corrected shot by shot, and are edited for free.
The real ceilings: 30 seconds on Wan 3.0, Wan 3.0 Prime and Seedance 2.5, 20 on Flux 3 Video, 15 on Kling V3, Seedance 2.0, Wan 2.7 and Grok, 8 on Veo 3.1. No model will give you a minute in one go, and that is just fine.
The method that works: cut it up, validate as images, generate five second shots, hold coherence with a Reference node and a Style node, chain with end frames when continuity demands it, then edit and score in one pass. Since editing is free, the only cost is the shots.
For a first attempt, take your first AI video if you have never generated one, then come back and make a thirty second film in six shots. Budget around 156 credits on Kling 2.5 Turbo, roughly one euro fifty, editing included.
Balance