August 12, 2026
Writing a prompt that works: the method that changes your images
Describe the photo rather than the idea, structure subject, action, setting, light, and style, iterate instead of piling on words: the complete method for writing effective image prompts on Imaginode, plus the honest list of what a prompt will never fix.
Describe the photo, not the idea
Here is the sentence that sums up this entire article: an image model does not understand your intentions, it executes your descriptions. "An image that inspires trust for my brand" is not a prompt, it is a brief for a human. The model knows neither your brand nor your definition of trust. It needs concrete material: objects, textures, a light.
Try the following mental exercise. Imagine the perfect photo already exists and you are describing it over the phone to someone who has to find it among a thousand others. You would not say "a warm image". You would say "a woman in her sixties laughs in a bright kitchen, hands covered in flour, blue apron". That is exactly the conversion you need to make.
That mental shift, from the idea to the photo, produces more progress on its own than any list of magic words gathered from social feeds. The "secret prompts" sold here and there are nothing but well-built descriptions. The good news is that the construction can be learned in an hour of practice.
The five-block structure that always holds up
A solid image prompt answers five questions, in this order: who or what, doing what, where, under what light, in what rendering style. Subject, action, setting, light, style. It is not a sacred formula, it is a checklist: every missing block will be filled in by the model with its statistical reflexes, in other words with generic filler.
A complete example: "a baker in his sixties, pulls a batch of baguettes from the oven, old-fashioned bakehouse with stone walls, warm early-morning light coming from a side window, documentary photography, 35 mm". Each segment answers one question, and none contradicts another. This prompt will produce variations, but all of them in the right direction.
Compare that with "authentic baker tradition artisan bread quality". No scene, no angle, no light: the model will receive those words as a vague mood and deliver the most average possible illustration of the bakery concept. Both prompts cost exactly the same price. The structure is free; you might as well take it.
The subject: a casting call, not a category
"A woman" is a category. "A businesswoman in her forties, short gray hair, burgundy suit" is a casting call. The difference comes down to a handful of visible details: an approximate age, one or two physical traits, a specific garment. Three choices are enough; the goal is not to draft a full identity record.
Choose the details that carry meaning for your image. If you are preparing a visual for an article about craftsmanship, the hands matter more than the hairstyle: "hands marked by work, short nails" will steer the model far more usefully than an eye color nobody will see in a wide shot anyway.
Same logic for objects and places. "A car" gives you a generic sedan with no identifiable make. "A family station wagon from the 90s, faded green paint, muddy plates" already tells a story. The model loves this kind of material: the more embodied the subject, the less the image smells like stock photography.
The action, or why your images look frozen
Many prompts describe a subject but forget to give it something to do. The predictable result: a character standing, facing the camera, with a passport-photo smile. If all your images have that catalog look, the missing action block is almost always the culprit. One concrete verb is enough to unlock the pose.
"Pours a coffee while looking out the window", "fixes a bicycle chain, brow furrowed", "runs through the rain holding her jacket over her head". The action gives the model a reason to place the arms, direct the gaze, create tension in the body. It generates composition naturally, without you having to describe any of it.
Be careful, though, with complex actions involving several people in precise interaction. "Two fencers, one parrying an attack in quarte" goes beyond what current models handle cleanly: you will get improbable limbs and fused swords. Stick to readable actions with a single main subject, at least while you are iterating at 1 credit.
The setting plants a mood, not an inventory
The setting places the scene and installs the atmosphere. Two or three well-chosen elements do the job: "restaurant kitchen in mid-service, stainless steel, steam" is enough to summon the entire world. No need to list the twenty expected utensils; the model knows restaurant kitchens better than you will ever manage to describe them.
The opposite error to the empty setting is the exhaustive one. A prompt that enumerates fifteen objects forces the model to cram them all in, and it will do it badly: floating objects, absurd scales, an overloaded composition. If a prop really matters, give it a place in the sentence; otherwise leave it to the model. You are describing a photo, not a stage play with its full prop list.
Think about depth too. "In the blurred background, shelves of books" places the element at the right distance and suggests a shallow depth of field in passing. These spatial cues, foreground, background, through a window, structure the composition far more than mood adjectives ever will.
Light, the most underrated parameter
Ask a photographer what makes an image: they will answer light before subject. Image models work the same way. The same sidewalk cafe becomes melancholic under "gray light of a rainy morning" and sun-drenched under "golden late-afternoon light, long shadows". A single segment of a sentence, and the entire emotion of the image tips over.
A few directions that work well: the golden light just before sunset, backlighting that cuts out a silhouette, the soft light of a north-facing window for portraits, colored neon for nighttime urban moods. Also specify where the light comes from when it matters: from the side, it sculpts faces; head-on, it flattens them.
If you were to add only one block to your current prompts, it would be this one. In our tests, specifying the light visibly improves the image more often than any other addition, notably because it unifies the entire scene. An average setting under beautiful light always makes a better image than the other way around.
Style closes the prompt and frames the rendering
Last block: telling the model what kind of image it is producing. "Realistic photography", "watercolor illustration", "minimalist 3D render", "60s screen-printed poster". Without that precision, the model chooses on its own, and its default choice varies from one model to the next. This is the block that prevents the surprise watercolor on a photography project.
For photography, a few technical terms pay off immediately: a focal length, 35 mm for a reportage feel, 85 mm for a portrait that separates the subject, an aperture like f/1.8 for a melted background. You do not need to understand optics; these terms work as shortcuts to renderings the models know by heart.
On Imaginode, the Camera node industrializes this part: five illustrated dials, framing, lens in millimeters, aperture, angle, and movement, which write the right English terms at the head of your prompt. To learn, turn the dials and watch the text they produce. It is a photography vocabulary course disguised as an interface.
Why English gets better results
Let's talk frankly about language. Image models accept French and manage fine with it, but their training data is overwhelmingly described in English. The direct consequence: English photographic vocabulary, golden hour, rim light, shallow depth of field, reaches far richer regions of their training than our French equivalents, when those equivalents even exist.
The gap is not cosmetic. On the same concept tested in both languages, the English version regularly produces more intentional framing and better-controlled light, simply because the precise terms exist and the model has seen them millions of times. "Contre-jour" is understood; backlit silhouette is mastered.
So should you learn photographic English before creating anything? No, and that is the whole point of the next section. But understanding why English wins will keep you from wrongly concluding that "AI does not work" when only the phrasing is limiting your results.
The magic wand, your technical translator at 1 credit
On every prompt field of Imaginode's Image and Video nodes, a magic wand rewrites your text into rich, structured English via the Kimi model, for 1 credit, about a cent. You write "a baker takes his bread out of the oven in the morning", and it develops the scene with the expected technical vocabulary: light, textures, rendering.
The right workflow: first write your complete idea in your own words, with your five blocks in place. Run the wand. Read what it produces, generate, then edit the English text directly if a detail drifts. The wand is not a lottery to rerun over and over; it is a translation and enrichment step, once per new idea.
A detail that matters for what comes next: the wand preserves the @mentions of your references. If your prompt contains @marie, your recurring character defined in a Reference node, the rewrite keeps the mention intact, with its description injected and its photos attached. Your enriched prompts therefore stay compatible with the entire reference system.
Iterate instead of piling on words
Faced with a disappointing image, the natural reflex is to add words. After five rounds of back and forth, the prompt is eighty words long, contradicts itself twice, and nobody knows anymore which segment produces what. This mille-feuille prompt is the most common dead end among intermediate creators, far more than the prompt that is too short.
The healthy method fits in one rule: one change per generation. The image is too dark? Modify only the light block, regenerate, compare. The framing disappoints? Only the framing. At 1 credit per attempt on Flux Schnell, ten disciplined iterations cost ten cents and teach you what every word weighs. It is a scientific experiment at a friendly price.
The node's history makes this discipline comfortable: the last 12 generations remain visible, each one comparable with the others. You literally see the effect of every word change by putting two versions side by side. When an iteration degrades the image, go back to the previous text and take another direction. Nothing is lost, everything is comparable.
Negative prompts don't exist everywhere
On forums you will run into prompts stuffed with "no blur, no extra fingers, no text". That is a legacy of older tools that offered a separate "negative prompt" field. On the recent models available in Imaginode, this syntax is not a standard: depending on the model, negation is understood unevenly, and merely mentioning a concept can be enough to summon it.
Write what you want to see, not what you want to avoid. Rather than "no crowd in the background", write "deserted street in the early morning". Rather than "not blurry", specify "sharp focus on the face". Positive phrasing describes a possible image; negative phrasing describes a void the model does not always know how to paint.
That leaves the defects nobody asked for, strange hands, unreadable text on signs. No negative incantation eliminates them reliably; it is a current limitation of the models, acknowledged by everyone. The realistic counter: iterate at low cost until you land a clean generation, or frame your scene so that hands and text are not in the foreground.
Match the length of the prompt to the model
Not every model digests the same density of text. Flux Schnell, at 1 credit, shines on short, direct prompts: a clear scene, a few specifics, and it delivers in seconds. Serving it a hundred-word slab improves nothing; it will skim half of it. It is the sketching tool; treat it as such.
High-end models like Flux 2 Pro at 7 credits or Seedream 5 Lite at 5 credits make far better use of detailed prompts: textures, complex light, nuanced mood. This is where text enriched by the magic wand earns its keep. The same generous description that drowned Flux Schnell becomes fuel for them.
Hence a two-phase method that echoes the golden rule from the beginner's guide: rough out short and cheap, finalize long and precise. Write the short version, iterate at 1 credit until the composition is right, enrich with the wand, then generate the final version on a 5 or 7 credit model. Your prompt grows along with your image.
What a prompt will never fix
It has to be said without detours: some problems cannot be solved in the text field, and grinding away at them costs whole evenings. The first is character consistency. You can describe your heroine across twenty lines, age, face, beauty mark: without reference images, she will change faces with every generation. That is not a flaw in your prompt, it is how the models fundamentally work.
The answer lives elsewhere on the canvas: the Reference node. You give it a name, a description, and photos, then you mention @thatname in any prompt. The mention shows in green, the description is injected, and the photos travel with the request automatically. Models like Seedream 5 Lite, the reference champion at 5 credits, lean on them to keep a face stable from one image to the next.
The same lucidity applies to the rest: long text rendered perfectly readable inside the image, a guaranteed five-fingered hand, a trademark reproduced faithfully, no prompt promises any of it. Recognizing these limits saves you considerable time: you will know when to refine the text, and when to switch tools. Character consistency deserves its own dedicated article on this blog, by the way.
A complete before and after to tie it all together
A real situation: a creator is preparing the visual for an article about remote work. First prompt, "woman remote work happy productivity". Predictable result: generic white desk, frozen smile aimed at the laptop, a stock photo without the license. The prompt described a concept; the model delivered the average cliche of that concept.
Second version, with the structure: "a woman in her thirties in a mustard sweater, focused, annotates documents in front of a laptop, apartment living room with green plants, soft late-morning light from a large window on the left, realistic photography, 50 mm, shallow depth of field". Enriched with the wand for 1 credit, iterated four times on Flux Schnell, finalized on Flux 2 Pro.
The tally: 12 credits, about twelve cents, a quarter of an hour on the clock. The final image has believable light, a natural pose, a real point of view. No hidden talent, no magic word: five blocks filled in honestly, a tool-assisted translation, and disciplined iterations. The entire method fits inside this one example.
Your next step
Take the last prompt you wrote, any prompt at all. Run it through the five questions: embodied subject, concrete action, situated setting, described light, specified style. Fill in the missing blocks, generate the new version at 1 credit on Flux Schnell, and compare the two images in the node's history. The gap should convince you better than this article ever could.
Then, two directions depending on your project. If your images feature a recurring character, the Reference node and @mentions are your absolute priority, well ahead of any vocabulary refinement. If cinematic rendering is what draws you in, explore the Camera node and its five dials, then our article on camera direction in AI.
And if you get stuck on a specific prompt, ask the assistant in the bubble at the bottom right: it sees your open canvas, your node, and your text, and answers for 1 credit. Describing the photo rather than the idea remains your job. For everything else, you are now equipped.
