TL;DR
- Problem: a 27B vision model read handwritten accident reports with 58% of critical fields right. Licence numbers were right 12 times out of 40.
- Finding: three prompt rewrites moved nothing. The model could barely see the digits.
- Fix: cut each field out at its known place and enlarge it 3 times. Critical fields went to 83%, licence numbers to 27 of 40.
- Scope: small handwriting on fixed forms. It did not help long free text.
One finding from the Constat OCR project, a pipeline built in 10 measured versions.
What I did
The form is a fixed template, so every field sits at a known place. That is all the method needs.
Photo
↓ align on the template
Page on the template
↓ cut each field
Crops
↓ enlarge 3x
Enlarged crops
↓ read all of them in one request
One JSON answer
↓ merge into the record
Record
| Step | What it does | Why |
|---|---|---|
| 1. Align | Warp the photo onto the blank form (ECC homography on the printed lines) | Known positions land on the right pixels |
| 2. Cut | Take each field’s box from the template plus a small margin | Handwriting runs a little outside the box |
| 3. Enlarge | Resize each crop 3 times | A digit gets several vision tokens instead of half of one |
| 4. Read | Send all crops of a form in one request, each labelled, and ask for JSON | One call per form, with the schema enforced |
| 5. Merge | A value that is read replaces the page value; null keeps it | A field is never lost |
| 6. Protect | Retry 3 times, never save a failed call | A network error cannot pass as a result |
Prompt tuning moved 0 points
Same 5 dev forms, 95 critical fields, three rounds:
| Round | Prompt change | Critical fields right (of 95) |
|---|---|---|
| 1 | Copy exactly, null over guessing, formats for dates, plates, numbers | 52 |
| 2 | + capitals for names, “the slash is thin” | 50 |
| 3 | + “the 3rd character is always the slash” | 49 |
Differences of 1 to 3 fields are noise: one digit flips between runs. The licence number stayed wrong on 9 of 10 vehicles, because the thin / was read as a 1 (631063647 for 63/068647).
The model could barely see the digits
The cause was not the words. It was the size.
- A vision model reads an image in patches of about 32 by 32 pixels, about one token each. A 4-megapixel page cost about 4,000 tokens.
- A handwritten digit on that page is about 15 to 20 pixels tall, so smaller than one patch.
- After a 3 times enlargement the same digit spans several patches.

The rule I took from it:
- Does the same error survive prompt changes?
- Compare the size of the detail with the model’s patch size.
- If the detail is smaller than a patch, enlarge it: crop, tile or upscale.
- Re-measure on the same set.
Enlarging moved 25 points
![]()
On the 20 test forms, with no prompt change and no second model:
| Critical fields right (of 380) | |
|---|---|
| Whole page | 222 (58.4%) |
| 6 digit fields enlarged | 299 (78.7%) |
| 22 fields enlarged | 315 (82.9%) |
| Field (of 40) | Whole page | Enlarged |
|---|---|---|
| Licence number | 12 | 27 |
| Attestation number | 16 | 30 |
| Policy number | 18 | 31 |
| Validity start date | 16 | 34 |
| Validity end date | 28 | 40 |
Four decisions I made from data
| Decision | Evidence | What I did |
|---|---|---|
| Stop prompt tuning | 3 rounds within noise | Changed the input instead |
| Don’t build a second-model vote | Two models agree on 52.7% of fields, and a rule to pick between them fixed 2 fields out of 980 | Didn’t build it |
| Drop 5 fields from the crops | Worse than the page read on the dev forms | Left them out |
| Discard a run | 8 of 20 forms silently fell back after network errors, so the score looked worse | Added retries, reran |
When it works, and when it doesn’t
| Works | Doesn’t |
|---|---|
| Fixed layouts, where field positions are known | Long free text: addresses, first name, place and prefecture got worse and were dropped |
| Small, dense handwriting | Genuinely ambiguous handwriting |
| Long digit strings | Layouts that move: you would need a field detector first |

The T in 14472-T-1 is written like a J. Enlarging does not make that unambiguous.
Limits and cost
- Synthetic forms: real ones hold personal data. The chart above uses 5 dev forms, and the 20 test forms point the same way.
- No human baseline: I did not measure how well a person reads these crops.
- Template positions: the crop positions come from the generator’s template. Real photos would need a field detector.
- Cost: about $0.01 of extra GPU per form for the crops, next to the page read.
Full version-by-version results: the Constat OCR project. Code: github.com/BENHIMA-Mohamed-Amine/constat-ocr.