September 5, 2026
Mangled hands and unreadable text: why AI fails at these two things and how to get around it
The real technical reason behind six fingers and invented words, the models that write correctly today, and the framing and retouching method that solves most cases without firing off ten more generations.

Frank HoubreFounder of Imaginode
The two flaws that give an AI image away in one second
Hands and text fail for opposite reasons: hands are too variable to be learned cleanly, text demands a symbolic precision the model never practises.
You show someone your image. It looks good, the light is right, the character stands up straight. And the person looks at the hand, counts the fingers, then reads the shop sign in the background where it says something like BAKERV.
These two flaws are the most recognisable signatures of image generation, and they have become the test everybody applies. What is frustrating is that they sit alongside otherwise excellent images: a portrait with perfect grain, ruined by a six fingered hand.
The good news is that the two problems have different causes and therefore different fixes. And that in 2026, one of the two is largely solved if you pick the right model. We are going to look at why it fails, then exactly what to do.
Why hands fail: what the model actually learned
A hand has more than twenty joints, shows up tiny and partly hidden in most training photos, and takes wildly different shapes depending on the angle: the model learned a blurry average of it.
A face is two eyes, a nose, a mouth, always in the same order, almost always front on or three quarters, and often the largest element in the photo. A model sees that hundreds of millions of times under comparable conditions. It learns very well.
A hand is the opposite on every count. More than twenty joints, so the number of possible positions explodes. It is small in the frame, often blurred because it moves, regularly hidden behind an object, a piece of clothing or the other hand. And depending on the angle, one and the same hand gives silhouettes that have nothing to do with each other.
The result is that the model outputs the average of hands, the same way it outputs the average of dragons when you ask for a dragon. Except an average dragon stays believable, nobody has seen one, whereas an average hand jumps out at anybody. And nothing in how it works tells it that one extra finger is a mistake: it has no rules, it has probabilities.
Why text fails: the model draws, it does not spell
An image model produces plausible shapes pixel by pixel, with no representation of letters as symbols: so it makes something that looks like writing rather than writing.
Text is a problem of a different nature. An image model does not know the alphabet. It does not know that a B is a B and that it comes before a C. It has seen billions of images containing areas that look like text, and it learned to produce areas that look like text.
That is why the failures have that characteristic look: it is never complete gibberish, it is almost right. The right proportions, the right letter spacing, the right typeface, and two invented letters in the middle. The model perfectly succeeded at what it was really doing, which was drawing a typographic texture.
Add that in an image, text is often tiny compared to the rest. The model allocates its attention to what fills the frame. A shop sign twenty pixels high does not get the care it would need, and that is where BAKERY loses a letter along the way.
What has actually changed on text
Recent models trained with strong language understanding write short text correctly and reliably: the text problem is today a question of which model you pick.
Here is the point a lot of people have not updated, because they formed an opinion in 2023 and never reopened it. Text in images works now, provided you bring out the right model.
Models that are language models first and draw afterwards have a structural advantage: they know what a word is before they draw it. Nano Banana Pro is built on Gemini 3 Pro and handles text in the image and precise layout. GPT Image 2 has exactly the same profile: its strength is not texture, it is obedience, and it writes short text cleanly.
Alongside them, Recraft V4.1 comes from the design world and slots short text cleanly into a graphic composition, which makes it the first reflex for a poster or a logo. And Flux 2 Max holds text a notch above Flux 2 Pro when you need photorealism with a readable sign in it.
The models not to use when there is text
Fast, cheap models produce very good images without knowing how to write: Flux Schnell at 1 credit stays unbeatable for exploring, unusable for a visual that carries a word.
The corollary of the previous paragraph, and the one that burns the most credits. Flux Schnell at 1 credit is perfect for exploring a composition, testing a framing, validating a mood. It will not write your tagline correctly, and firing it off fifteen more times will not change that.
The expensive reflex is to insist with the model you started on. You are on your twelfth generation, you have spent twelve credits for nothing, when a switch to a model that knows how to write would have fixed it first time for 8 or 19 credits.
Never fire the same model more than three times at a text flaw. If three tries do not get there, it means the model cannot do it, not that you asked badly. Change model, that is all.
How to write the instruction so the text comes out right
Put the exact text in quotation marks, explicitly require that it appear spelled that way and that no other word appear, and keep it to a few words.
The wording counts as much as the model. Write the exact text in quotation marks in your prompt, then add a sentence that locks it down: must appear, spelled exactly like this, perfectly readable, no other word. Without that sentence, the model will happily add a strapline of its own invention under your title.
Second rule: stay short. A brand name, a five word tagline, a date. A whole paragraph inside an image, no model will do that cleanly today, and it is not a good idea graphically either. Long text goes on in editing or in a layout program, not in the generation.
Third rule: give it room. Explicitly ask for the text to occupy a large, well contrasted area. A model asked to write six letters on a sign at the end of the street will fail where it succeeds on a full frame title. It is the same logic as for any instruction, and the subject is covered more broadly in writing a prompt that works.
Hands: three framing choices that solve most of it
Frame above the hands, keep them busy with an object, or move them out of the sharp plane: these three decisions eliminate most failures without costing a credit.
On hands, nobody has a miracle solution and I am not going to invent one for you. Recent models handle them noticeably better than two years ago, but none guarantees five fingers every time. What works is not asking them for the hardest exercise.
First choice: the framing. A chest shot or a face close up removes the problem by construction. Look at professional advertising photography, it frames above the hands far more often than people think. The Camera node lets you set that framing cleanly instead of hoping for it.
Second choice: keep the hand busy. A hand holding a cup, a phone, a pen, a steering wheel, it does not matter, is a hand whose position is constrained by the object, and models handle those configurations far better than an open hand in mid air. Third choice: take it out of the sharp zone. An f/1.4 aperture in the Camera node puts the hands in the background blur, and a blurred finger does not get counted.
Describe the action rather than the anatomy
Counting fingers in the prompt does not work: the model does not count, you have to describe the gesture, what the hand is holding and how it holds it.
The universal reflex is to write five fingers in the prompt. It does not work, and it will not work, because the model has no counting mechanism at all. You can write exactly five fingers, an anatomically correct hand, it has no rule to apply.
What works is describing the scene so the position is implicit. Both hands flat on the table. One hand closed around a cup handle. Arms crossed. Hands in pockets. You are describing a configuration the model has seen millions of times in a stable position, and it renders it back.
It is the same principle as for everything else, really: you get far more by describing a concrete situation than by giving the model abstract instructions. Describing an action, it can do. Following a numeric constraint, it cannot.
Repair instead of regenerating: targeted editing
An editing model takes your image and only changes what you describe: fixing a hand costs a few credits against a whole new generation and a new framing to validate.
Here is the habit that changes everything, and the one beginners do not have. When an image is 95 percent good, you do not regenerate. You repair.
Flux Kontext at 6 credits never starts from scratch: it requires an input image and transforms it on instruction, preserving what pure generation models would recompose at random. You give it your image and describe only what has to change: fix the left hand, it should hold the cup by the handle, change nothing else.
Nano Banana 2 at 10 credits does the same job with an understanding of compound instructions, and GPT Image Mini at 2 credits is the cheapest option in the catalogue for a simple retouch. In every case, the second half of your instruction must list what does not move: face, pose, framing, light, set. Without that list, the model hands you back another image, a very good one, but another one.
The zoom before delivery, the proofread everybody skips
Flaws show at one hundred percent, not on a thumbnail: enlarging the image before delivering it costs 12 credits and avoids the mistake you discover at print time.
A generated image gets looked at on a canvas thumbnail, approved, then sent. Three weeks later it is printed at A2 and the hand has six fingers. It happened to me, and once is enough to pick up the habit.
The proofread takes two minutes. You open the image large and sweep three zones in order: the hands, everything containing text including backgrounds, then the eyes and the teeth. Those three passes catch nearly every flaw.
For a visual destined for print, the Enhance node with Topaz Precision at 12 credits enlarges cleanly and reveals along the way what the low resolution was hiding. It is the same tool as the one described in making a photo sharp, and it serves as an inspection magnifier as much as an enlarger.
The other details that give an AI image away
Teeth too regular, jewellery that doubles up, incoherent fabric patterns, reflections that match nothing: the same causes produce other flaws, less famous but just as visible.
Hands and text get top billing, but the same mechanics produce other failures. Teeth, often too numerous and too regular, because they are small and repetitive. Jewellery and glasses, which double up at the arms. Fabric patterns, stripes or checks, which head in one direction then another in the middle of the garment.
Reflections are a special and very telling case: in a mirror or a shop window, the model draws something that looks like a reflection without checking that it matches the scene. Look at the glass behind your characters, you will regularly find things there that exist nowhere in the image.
And crowds, where the faces in the background turn disturbing as soon as you zoom in. There again the reflex is the same: a framing that avoids the problem, or targeted editing that fixes it. Those are exactly the points I review in beginner mistakes in generative AI.
The procedure, in short
Pick a model that writes when there is text, frame to avoid open hands, repair by editing rather than regenerating, and proofread large before delivering.
On text, the problem is solved and it is a question of model: Nano Banana Pro, GPT Image 2 or Recraft for graphic work, Flux 2 Max when you need photorealism with text in it. Exact text in quotation marks, locking sentence, a few words maximum.
On hands, no model is perfect, and the answer is in the staging: frame above, keep the hand busy with an object, or take it out of the sharp plane. Describe the gesture, never the number of fingers. And when an image is 95 percent good, repair it with an editing model instead of starting over.
To test all this without burning a budget, explore at 1 credit on Flux Schnell and only move to an 8 or 19 credit model on the composition you keep. The exact price shows on the button before every click, and the full grid is here.