September 5, 2026
Should you write your prompts in English? The answer depends on the model
What actually reaches the model when you type in your own language, why English was compulsory for so long, the models that now understand other languages without loss, and the 6 credit test that settles it for your case.

Frank HoubreFounder of Imaginode
The question everybody has without daring to ask it
Your prompt reaches the model exactly as you typed it, with no translation in between: the language you use therefore has a real effect on the result.
It is one of the first questions people ask, often in a low voice as if it were a silly question. It is not silly at all. It has a real technical answer, and that answer has changed recently.
Let us start with the most important point, the one nobody mentions: there is no automatic translation between your keyboard and the model. What you type goes off as is. If you write in French, French is what arrives at the model, with everything that implies.
So, your language or English? The honest answer is that it depends on the model you picked, and that a three minute test settles it for your precise case. We are going to do both: understand why, then test.
Why English was the only right answer for so long
Image models learned from images captioned overwhelmingly in English: a word in another language was under represented there, so less firmly attached to what it describes.
An image model learns by looking at hundreds of millions of image and caption pairs. Those captions come from the web, and the web indexed for that kind of task is massively English speaking.
Mechanical result: the phrase golden hour was seen millions of times stuck to photos taken at the end of the day. Its equivalent in French, heure dorée, far less. The model therefore had a much sturdier association for one than for the other, and a French prompt often gave a vaguer, more generic, less controlled render.
That is where the universal advice prompt in English comes from, repeated everywhere since 2022. It was right. It is far less right today, and repeating it without nuance makes people miss something important.
What changed: the models that understand before they draw
Recent models backed by a large language model interpret the instruction before generating, and that understanding is genuinely multilingual.
The new generation of image models no longer just maps words to textures. Some of them go through a language model that reads and understands your request, then steers the generation.
Nano Banana Pro is built on Gemini 3 Pro and is the most capable in the catalogue at understanding long instructions. Nano Banana 2 runs the image generation of Gemini 3.1 Flash. GPT Image 2 comes from the same world, and its strength is explicitly obedience to the instruction rather than texture.
Those models understand French the way they understand English, because the part that understands is a language model trained in dozens of languages. On them, writing in your own language costs you next to nothing, and you gain precision since you are writing in the language you think in.
The models where English is still preferable
Pure image models, fast and cheap, stay more faithful in English: it shows most clearly on technical vocabulary.
At the other end of the catalogue, classic diffusion models have no linguistic understanding layer. They encode your text and use it as a statistical guide. There, the language still counts.
Flux Schnell at 1 credit, the Flux family in general, and anything running locally on Stable Diffusion fall into that category. You will get more faithful results in English, especially as soon as the prompt contains precise vocabulary.
This is not a question of model quality, by the way: Flux 2 Pro makes very beautiful images. It is a question of how it reads your text. An excellent model can be a poor reader of French.
The 6 credit test that settles it for your case
Generate the same prompt three times in your language, then three times in its English translation on the model you use: the gap shows up immediately or does not exist.
Rather than taking my word for it, run the test. It costs 6 credits on Flux Schnell and takes three minutes, and it holds for your model, your subject, your style.
Write a prompt containing slightly technical vocabulary: a low angle shot of a woman at the edge of a cliff, end of day light from behind, shallow depth of field. Generate three times. Then translate the same prompt into English and generate three times on the same model, same format.
Look at the six images together. If the three English ones hold the low angle and the backlight better, your model prefers English and you will know it once and for all. If the gap is not visible, write in your own language with no qualms: you will be more precise in it, and precision counts for more than language.
The vocabulary where English keeps a real advantage
Technical cinema and photography terms are very dense keywords in English, whose equivalent in other languages is often vaguer or longer.
Even on a model that understands French perfectly, certain English words stay superior, and it is a question of density rather than of language.
Take bokeh, golden hour, backlit, rim light, low angle, shallow depth of field, dolly in. Each of those terms designates one precise thing and one only, and the model has seen them millions of times attached to exactly that effect. The French equivalent often needs a paraphrase, and a paraphrase dilutes.
Hence the mixed method, which is what I do in practice: the scene, the action and the intent in my own language, the technical terms in English. A golden hour backlight on a running woman, shallow depth of field. It is ugly to read and it works very well, because a prompt is not a text, it is a list of instructions.
The Camera node already writes the English for you
The Camera node settings reach the model as English prompt fragments: the language question is already settled for everything to do with staging.
There is a reason I barely write technical vocabulary in my prompts any more, and it is structural. The Camera node contains menus for framing, lens, aperture, angle and movement.
What you do not see is that every option in those menus is stored as an English prompt fragment, and it is that fragment which goes to the model. You pick close up in an interface in your own language, and the model receives the English term it knows by heart.
The benefit goes beyond translation. You no longer have to memorise the vocabulary, you no longer get a term wrong, and the setting plugs back into several nodes to keep things coherent. The detail of what each setting changes is in directing the camera in AI.
Text inside the image: a completely different question
The language of the prompt and the language of the text to display are independent: you get a French word by giving it in quotation marks, whatever language the instruction is in.
Careful not to mix up two things. The language you give your instructions in has nothing to do with the language of the text that must appear in the image.
If your poster has to say BOULANGERIE, write the exact word in quotation marks in your prompt, adding that it must appear spelled exactly like that and that no other word must appear. The rest of the instruction can be in English, it changes nothing.
Two extra precautions for French. Accents come through less reliably than unaccented letters on some models, so check the é and the à at full size. And take a model that knows how to write: Nano Banana Pro, GPT Image 2 or Recraft V4.1. The full subject is in mangled hands and unreadable text.
The false friends that ruin a generation
Translating a technical term word for word regularly produces the opposite of the intended effect: a few common words are known traps.
If you switch to English, do not translate with a dictionary. A few examples that come up all the time and cost generations.
Plan, in French, means a framing or a take. Translated as plan in English, it becomes a project or a diagram. The right word is shot. Cadre translates as frame, not as box. Ambiance gives mood or atmosphere, never ambiance on its own, which exists but means something else in everyday English. And net is sharp or in focus, certainly not net.
The safe reflex is to have a language model write the English prompt rather than translating it yourself. The Text node does it for 1 to 3 credits: give it your idea in your language and ask for an image prompt in English, detailed, cinema vocabulary, no title and no commentary.
And for video, is it the same?
Video models follow the same logic, but the video prompt is so short when you start from a still image that the question loses almost all its interest.
The rule is the same: video models backed by strong linguistic understanding take French well, the others prefer English. Nothing new on that front.
Except that in practice, the question comes up far less. When you work cleanly, the Video node receives an already validated starting image, and the prompt only describes movement. The camera moves slowly forward. The hand closes. Three words, where the language no longer changes much.
It is a pleasant side effect of the storyboard method, described in making your storyboard in AI before paying for the video: by shifting the load from text to image, you make the language question almost moot.
What I actually do
My own language for the scene and the intent, English for the technical terms, the Camera node for staging, and a six image test whenever I discover a model.
My practice fits in three lines. I write the scene, the action and the intent in my own language, because I am more precise in it and precision is what counts most in a prompt. I keep the technical terms in English, because they are denser.
Everything to do with staging goes through the Camera node, so it is no longer in my text at all. And when a new model enters the catalogue, I run the six image test before forming an opinion. It costs 6 credits and saves months of false belief.
The advice I no longer give is prompt in English without thinking. That was true three years ago, today it loses precision on half the catalogue. And a precise prompt in your own language always beats an approximate one in English.
In short
On models backed by a language model, write in your own language; on pure diffusion models, switch to English or to a mix of the two.
Nano Banana 2 and Pro, GPT Image 2: French comes through with no notable loss, write in your language. Flux Schnell, the Flux family and everything running locally: English stays more faithful, especially on technical vocabulary.
In every case, cinema vocabulary in English remains a good reflex, and the Camera node spares you most of it. Text to display inside the image is given in quotation marks, independently of the language of the instruction.
And if you want to dig into the structure of a prompt rather than its language, that is where the real difference in results is played out: writing a prompt that works details the full anatomy, subject, action, framing, light, style.