← Back to Blogs

Cropping is all you need in OCR, not prompt tuning

Cropping is all you need in OCR, not prompt tuning

TL;DR

  • Problem: a 27B vision model read handwritten accident reports with 58% of critical fields right. Licence numbers were right 12 times out of 40.
  • Finding: three prompt rewrites moved nothing. The model could barely see the digits.
  • Fix: cut each field out at its known place and enlarge it 3 times. Critical fields went to 83%, licence numbers to 27 of 40.
  • Scope: small handwriting on fixed forms. It did not help long free text.

One finding from the Constat OCR project, a pipeline built in 10 measured versions.


What I did

The form is a fixed template, so every field sits at a known place. That is all the method needs.

Photo
  ↓ align on the template
Page on the template
  ↓ cut each field
Crops
  ↓ enlarge 3x
Enlarged crops
  ↓ read all of them in one request
One JSON answer
  ↓ merge into the record
Record
StepWhat it doesWhy
1. AlignWarp the photo onto the blank form (ECC homography on the printed lines)Known positions land on the right pixels
2. CutTake each field’s box from the template plus a small marginHandwriting runs a little outside the box
3. EnlargeResize each crop 3 timesA digit gets several vision tokens instead of half of one
4. ReadSend all crops of a form in one request, each labelled, and ask for JSONOne call per form, with the schema enforced
5. MergeA value that is read replaces the page value; null keeps itA field is never lost
6. ProtectRetry 3 times, never save a failed callA network error cannot pass as a result

Prompt tuning moved 0 points

Same 5 dev forms, 95 critical fields, three rounds:

RoundPrompt changeCritical fields right (of 95)
1Copy exactly, null over guessing, formats for dates, plates, numbers52
2+ capitals for names, “the slash is thin”50
3+ “the 3rd character is always the slash”49

Differences of 1 to 3 fields are noise: one digit flips between runs. The licence number stayed wrong on 9 of 10 vehicles, because the thin / was read as a 1 (631063647 for 63/068647).


The model could barely see the digits

The cause was not the words. It was the size.

  • A vision model reads an image in patches of about 32 by 32 pixels, about one token each. A 4-megapixel page cost about 4,000 tokens.
  • A handwritten digit on that page is about 15 to 20 pixels tall, so smaller than one patch.
  • After a 3 times enlargement the same digit spans several patches.

The same licence number field twice with the model's 32 pixel patches drawn in red: on the whole page a digit is smaller than one patch, after enlargement it spans several

The rule I took from it:

  1. Does the same error survive prompt changes?
  2. Compare the size of the detail with the model’s patch size.
  3. If the detail is smaller than a patch, enlarge it: crop, tile or upscale.
  4. Re-measure on the same set.

Enlarging moved 25 points

Critical fields right out of 95 on the same 5 dev forms: 52, 50 and 49 for three prompt rewrites, then 80 and 83 for enlarging 6 and 22 fields

On the 20 test forms, with no prompt change and no second model:

Critical fields right (of 380)
Whole page222 (58.4%)
6 digit fields enlarged299 (78.7%)
22 fields enlarged315 (82.9%)
Field (of 40)Whole pageEnlarged
Licence number1227
Attestation number1630
Policy number1831
Validity start date1634
Validity end date2840

Four decisions I made from data

DecisionEvidenceWhat I did
Stop prompt tuning3 rounds within noiseChanged the input instead
Don’t build a second-model voteTwo models agree on 52.7% of fields, and a rule to pick between them fixed 2 fields out of 980Didn’t build it
Drop 5 fields from the cropsWorse than the page read on the dev formsLeft them out
Discard a run8 of 20 forms silently fell back after network errors, so the score looked worseAdded retries, reran

When it works, and when it doesn’t

WorksDoesn’t
Fixed layouts, where field positions are knownLong free text: addresses, first name, place and prefecture got worse and were dropped
Small, dense handwritingGenuinely ambiguous handwriting
Long digit stringsLayouts that move: you would need a field detector first

Two enlarged crops of handwriting: a plate written like 1.447Z -J-1 whose T looks like a J, and a licence number

The T in 14472-T-1 is written like a J. Enlarging does not make that unambiguous.


Limits and cost

  • Synthetic forms: real ones hold personal data. The chart above uses 5 dev forms, and the 20 test forms point the same way.
  • No human baseline: I did not measure how well a person reads these crops.
  • Template positions: the crop positions come from the generator’s template. Real photos would need a field detector.
  • Cost: about $0.01 of extra GPU per form for the crops, next to the page read.

Full version-by-version results: the Constat OCR project. Code: github.com/BENHIMA-Mohamed-Amine/constat-ocr.