TL;DR
I built a pipeline that reads photos of handwritten Moroccan accident reports, and measured every step. Over 10 controlled versions the character error rate fell from 89% to 2.5%, about 36 times lower, and critical fields right rose from 0.8% to 82.9%, with no fine-tuning.
- The result: character error rate 89% to 2.5%. That is the share of characters that must be edited to match the truth, counted over every field of every form, so a wrong digit costs. 97% of the character errors are gone. Critical fields: 315 of 380 right (v1: 3). Minor fields: 501 of 600 (v1: 12). Ticks and categories exact. 4 of 20 forms with every critical field right (v1: 0). Exact-match scoring on a frozen benchmark, no LLM judge.
- The stack: one open-weight 27B vision model (Qwen3.8-27B, Apache 2.0) on a Modal H100 behind a proxy token. No OCR engine, no text LLM, no third-party LLM API: the form never leaves my own deployment.
- Three moves did the work. I found the real bottleneck with data (92% of v1’s errors were values the OCR never read). I read ticks with no model at all, because the form is a fixed template. And when prompt tuning stalled, I changed the pixels: crop each field at its known place and enlarge it, which took critical fields from 58% to 83%.
- The discipline: one variable per version, harness before model, the bad numbers published (v1 scored 0.8% on purpose). When a run silently skipped 8 of 20 forms on network errors, I threw it out, fixed the cause and reran it.
- The honest limits: synthetic forms (real ones hold personal data), 20 test forms for the pipeline headline, and a person is still needed on 16 of 20 forms. About $0.04 of GPU per form.
Related: the blog post Cropping is all you need in OCR, not prompt tuning explains the v4b and v4c finding in 3 minutes.
| Version | The one change | Critical fields right (of 380) | Character error rate |
|---|---|---|---|
| v1 | Tesseract + LLM (baseline) | 3 (0.8%) | 89% |
| v2 | PP-OCRv6 reads the photo, on CPU | 56 (14.7%) | 70.1% |
| v3a | Straighten the photo first | 54 (14.2%) | 66.8% |
| v3b | Sort the text by zone of the form | 55 (14.5%) | 53.2% |
| v3c | A vision model that reads handwriting, on a GPU | 168 (44.2%) | 16.8% |
| v3d | Repair the record with deterministic rules | 194 (51.1%) | 16.1% |
| v3e | Read ticks, tiles and circles from the template, no model | 194 (51.1%), marks now exact | 16.1% |
| v4a | One open vision model reads the page: no OCR step, no LLM API | 222 (58.4%) | 5.3% |
| v4b | Re-read the 6 hardest digit fields from enlarged crops | 299 (78.7%) | 3.7% |
| v4c | Crop 22 fields, header included | 315 (82.9%) | 2.5% |
What this shows about how I work
- Evaluation before models. The frozen dataset, six metrics, tests and CI existed before v1 had a single result. Every later number is a fair comparison, not a guess.
- Diagnose, then fix. Each version exists because the one before it measurably failed at something specific. The error analysis said where to look: the reader, then the layout, then close misses, then pixels.
- Logic before models. OpenCV to straighten, position to sort the text, a template reader for ticks, small rules for the record. A model is used where nothing simpler works.
- Models tested, not trusted. Two famous OCR models topped public rankings and failed on the handwritten form. A general open VLM then beat the specialist chain. Neither result was in a leaderboard.
- Production thinking. Serverless GPUs that scale to zero, a proxy token that answers 401 before any GPU starts, pinned model and vLLM versions, retries, and a cost figure that counts the GPU, so versions are compared fairly.
- Honesty. I report the bad numbers, the invalid run, the rule that only holds on synthetic data, and what the scores do not show.
Reading the numbers: the metric is strict, on purpose
- Scoring is exact match and deterministic. There is no LLM judge, so the same run always gives the same number.
- A single stray character, a space, or a swapped pair of dates counts as a wrong field.
- Why the character error rate is the number I lead with: field accuracy only says whether a value is exactly right. The character error rate also says how far off the wrong ones are, so it shows progress that field accuracy hides. In v3b the critical fields barely moved (14.2% to 14.5%) while the character error rate fell 13 points, because values started landing in the right vehicle’s fields. In v4 the two moved together.
- The synthetic forms are neater than real ones (their ticks and circles are clean), so the comparison between versions matters more than the absolute score. This is not a state-of-the-art claim: it is a controlled study on my own benchmark.
Why this form
Filling in accident reports by hand is slow and repetitive, and for an insurer time is money, so it is a real automation target with a clear payoff. The input is a photo of a filled-in constat amiable, Morocco’s handwritten two-driver accident report. The output is the same data, structured. Anyone can call an LLM on an image and get something back. What matters in production is knowing whether the result is good enough to ship, and proving it stays good as the pipeline changes. So I treated OCR as a testing and evaluation problem first, not a model-picking problem.
- The form is hard on purpose: handwriting, French and Arabic labels, ticked boxes, a hand-drawn sketch, a circled licence category. Easy printed text would have taught me nothing.
- The harness came before any model: a frozen dataset, fixed metrics and CI regression tests on every push, all built before v1 had a result. That is what makes every later version a fair comparison.
- Real constats hold personal data, so the project generates its own: fake but internally consistent forms with handwriting, ticks, and a sketch whose two cars collide where the boxes say they should. The record that generated each form is its answer key, so accuracy is measured against ground truth, not judged by eye.
- Every version is judged on the same four things: quality, speed, cost, security.
v1: the naive baseline
How it works
photo → Tesseract (French OCR) → raw text → Groq gpt-oss-120b (JSON mode) → structured record
- Pipeline: Tesseract reads the photo and outputs raw text. Groq’s
gpt-oss-120bturns that text into a structured record, using JSON mode. The expected output format is generated directly from the project’s Pydantic schema, so the prompt can never drift out of sync with the actual data model. - Deliberately basic. No image cleanup, no vision model, plain text OCR only. The goal of v1 is an honest baseline and a clear list of what actually breaks, not a good score.
Results
Measured on 20 test forms (evaluation size was limited by Groq’s free-tier rate limit).
| Metric | Result |
|---|---|
| Critical fields correct | 3 of 380 (0.8%) |
| Forms with zero critical errors | 0 of 20 |
| Character error rate | 89% |
| Vehicle type / impact zone / licence category | 0 of 40 each, below even the “always guess the most common answer” baseline |
| Cost | $0.0009 per form |
| Speed | 23s per form on average. Most of that is waiting on Groq’s free-tier rate limit, not actual processing: in a smaller 2-form test run before the rate limit kicked in, Tesseract took about 2.3s and the Groq call about 2.5s, so roughly 5s of that time is real work |
- Not good, and that was expected. v1 is a naive baseline on purpose.
- The real finding: 92% of the wrong values were never present in Tesseract’s OCR text at all. The LLM only lost an extra 8% on top of that. The bottleneck is the OCR step, not the LLM.
What it showed
- OCR was the bottleneck. The LLM was not the problem, which is why v2 changed only the reader.
- Two production blockers surfaced: a free-tier rate limit that capped the evaluation at 20 forms, and data residency, because every document went to a third-party API. The rate limit is why later runs use my own GPU deployment, and v4a removed the third-party LLM API altogether.
Update: v2, CPU OCR
v1 showed the OCR step was the bottleneck, so v2 changes that one thing: PP-OCRv6 through RapidOCR (ONNX Runtime, CPU only) replaces Tesseract. Same gpt-oss-120b, same prompt, same 20 test forms. A new engine is a new class behind the existing OcrEngine interface, so no working code was rewritten.
- Chosen on my forms, not on a leaderboard. The shortlist came from a public benchmark (OmniDocBench), then four engines ran on 5 dev forms, with the test forms kept for the winner only. The benchmark did not predict the winner.
| Rank | OCR engine | Critical fields right (of 95) | Character error rate |
|---|---|---|---|
| 1 | PP-OCRv6 (RapidOCR) | 17 | 65.4% |
| 2 | PP-OCRv5 (RapidOCR) | 6 | 72.8% |
| 3 | docTR | 4 | 82.1% |
| 4 | Tesseract (v1) | 1 | 87.3% |
- Result on the 20 test forms: critical fields 0.8% to 14.7%, character error rate 89% to 70%, still no form fully right. 81% of the wrong values were never in the OCR text.
- Vehicle B was nearly unread (7 of 440 text fields right, against 105 for vehicle A), and ticks, vehicle type and impact zone stayed at zero: they are pictures, which no text-only pipeline can read.
Update: v3a, straighten the photo
v2 read the photo as it was: tilted, and seen at an angle. v3a flattens it first, and changes nothing else.
- What changed: an OpenCV step finds the page against the desk, takes its four corners and warps it flat. A new step in the pipeline, off by default so earlier runs stay reproducible.
- A first attempt was dropped: a ready-made scanner package turned out to download a 1 GB background-removal model. Plain OpenCV does the job with no model.
- Result: almost nothing. Critical fields 54 of 380 (v2: 56), character error rate 66.8% (70.1%).
- Finding: the OCR engine already coped with the tilt. The unread values were handwriting it misread, not geometry. The step stayed because the next one needs the page upright.
Update: v3b, sort the text by zone of the form
In v2 and v3a, vehicle B’s values were in the text but mixed with the other column, and in v3a the LLM answered null for 97 of them.
- What changed: each text box is placed in one of four zones (header, vehicle A, circumstances, vehicle B) by its position, and the LLM gets four labelled blocks instead of one mixed stream.
- Per-form cuts: the page shifts by about 2% between forms, so the cuts are found on each form from the two printed green strips, not fixed.
- A rejected idea, measured: reading each column as its own image found 10% more text lines but the same number of correct values, so the page is read once.
- Result: the headline barely moves (critical 55 of 380), but vehicle B text fields right went from 7 to 26 of 440 and the character error rate fell from 66.8% to 53.2%.
- Finding: layout was a real problem, but a small one. The ceiling is the reading: 676 of 980 values were never in the OCR text.
Update: v3c, a vision model that reads handwriting
v3b showed the reader was the limit, so v3c changes only the reader: an OCR vision-language model instead of a text OCR engine.
Models tried, not trusted
- PaddleOCR-VL and DeepSeek-OCR-2 top public rankings, but those score printed pages. On the handwritten form, both read the header and then looped or invented text until the token cap. Dropped, with the runs to show it.
- Chandra-OCR-2 (Datalab, 5.3B, handwriting capable) found 28 of 49 values on a test page against 18 for the previous engine, read vehicle B, and returned ticked boxes. It was chosen on our own forms, not on a leaderboard.
- A general VLM was deliberately not the first step. I kept it for later, and it did win: see v4a.
Serving the model: secure, scalable, reproducible
- Serverless GPUs: Modal, with vLLM behind an OpenAI-style API. One small deploy file per model.
- Secure: every request needs a Modal proxy token. With no token the proxy answers 401 before any GPU container starts, so a leaked URL cannot run up the bill. Tokens live in the environment, not in code.
- Scalable: the service scales to zero after 2 idle minutes, batches concurrent requests on one GPU, and the pipeline runs forms in parallel (13 pages at once in the test run).
- Fast cold starts: weights and compiled kernels are cached in volumes, so a cold start on the H100 dropped from about 190 s to about 95 s once the cache was warm.
- Reproducible: the model revision and the vLLM version are pinned.
- Understood, not guessed: generation speed is set by memory bandwidth (tokens per second is about bandwidth divided by weight size). An L4 predicted 28 tokens per second and measured 27.7. An H100 was about 8 times faster per token.
- Loops are handled: the same page finished on one GPU and looped on another, so the engine regenerates a looping answer at a higher temperature, as the model’s own client does.
Results (same 20 test forms)
| Metric | v3b | v3c |
|---|---|---|
| Critical fields right | 55 of 380 (14.5%) | 168 of 380 (44.2%) |
| Minor fields right | 99 of 600 (16.5%) | 377 of 600 (62.8%) |
| Character error rate | 53.2% | 16.8% |
| Vehicle B text fields right | 26 of 440 | 200 of 440 |
| Ticks found (of 25) | 0 | 10 |
| LLM cost per form | $0.0011 | $0.0012 (GPU time not included) |
- The reader was the ceiling, and it moved: values never in the OCR text fell from 676 to 377 of 980.
- Still weak: long digit strings (attestation, licence and policy numbers, plates), and the category fields (vehicle type, licence category, impact zone), which were worse than always guessing the most common value (fixed in v3e).
Update: v3d, repair the record with rules
An error analysis showed most wrong answers were close, and some were plain rules, not reading.
- The analysis: 57% of the wrong text fields were one or two characters off. Vehicle B’s validity line is written right to left, so its two dates came out swapped. ID numbers carried stray spaces.
- What changed: an optional step after the LLM applies small deterministic rules: dates put in order, phone and policy numbers without separators, the attestation number format. Each rule is a small class in a registry, and every change is reported.
- A clean comparison: the saved v3c outputs were scored again, so the rules are the only difference and no GPU or LLM call was needed.
- Result: critical fields 168 to 194 of 380 (51.1%). 40 fields changed, 30 became right and none that was right became wrong.
- Honest about one rule: the attestation format comes from the synthetic forms, so it is only trusted there. Without it the gain is +19 critical fields instead of +26.
- Left out on purpose: snapping names and insurers to lists, because the lists would come from the data generator and inflate the score without the pipeline improving.
Update: v3e, ticks and tiles from the template
After v3d the mark fields were still poor (ticks 10 found of 25, vehicle type 12 of 40, licence category 2 of 40). But the form is a fixed template, so a tick is not a reading problem, it is a measurement at a known place.
- What changed: an optional step after the LLM aligns the straightened page on the blank template and reads each mark with plain image processing: ink in each checkbox, colour in the vehicle-type tile, ink on a ring around the licence letters, the blue impact patch, ink in the OUI or NON cells. No model, no GPU, no LLM.
- The geometry comes from the template PDF itself: checkbox squares, picture frames and letter positions.
- What made it work: aligning the page once was not enough, because it is a few pixels off at one end of the column. Each box is placed on its own printed square. A first version missed 4 of the 130 dev ticks; after the fix the gap between the weakest tick and the strongest empty box went from negative to a clear margin.
- Thresholds on dev only: set on 100 dev forms, then the 400 held-out test forms were read once.
| The reader alone | Dev, 100 forms | Test, 400 held-out forms |
|---|---|---|
| Ticks found / missed / extra | 130 / 0 / 0 | 542 / 0 / 0 |
| Vehicle type, licence category, impact zone | 100% | 800 of 800 each |
| Other damage (OUI or NON) | 100% | 400 of 400 |
| In the pipeline (20 test forms) | v3d | v3e |
|---|---|---|
| Ticks found / missed / extra | 10 / 15 / 4 | 25 / 0 / 0 |
| Vehicle type | 12 of 40 | 40 of 40 |
| Licence category | 2 of 40 | 40 of 40 |
| Impact zone | 4 of 40 | 40 of 40 |
| Other damage | 8 of 20 | 20 of 20 |
-
A clean comparison: the saved v3d outputs were scored again, so the step is the only difference. The text metrics do not move, since the step reads marks, not text.
-
Honest about what it shows: the synthetic forms have clean marks (ink crosses, one circle, a colour fill, a blue patch), so this is a baseline for what the template alone can tell. Real photographed marks are messier. The tick count is the number of ticks, not the handwritten digit.
-
Where this went next: the text was still the weak part (405 wrong fields), so v4 went after it. The final numbers for the whole series are in the TL;DR above.
Update: v4a, one open vision model reads the page
v3e left 405 wrong text fields, and 44% of them had never been read. v4a removes the hop where values were lost: one general vision model sees the page and the schema together.
- What changed: Qwen3.8-27B (Alibaba, Apache 2.0) on a Modal H100 with vLLM reads the straightened photo and returns the record as JSON that follows the schema (guided decoding, so keys and types are always right). No OCR engine, no text LLM, no Groq call. The tick reader and the repair rules still run after it.
- Data residency, solved: v1 listed it as a blocker, because every document went to a third-party API. In v4 the form goes only to an open model on my own deployment.
- Prompt: about 20 lines. Copy exactly, use null rather than guess (a plausible value that is not on the page is worse than null), read each digit alone, and the written format of each value.
- Same server rules as before: proxy-token security, scale to zero, pinned model revision and vLLM version.
| Metric (same 20 test forms) | v3e | v4a |
|---|---|---|
| Critical fields right | 194 of 380 (51.1%) | 222 of 380 (58.4%) |
| Minor fields right | 381 of 600 (63.5%) | 482 of 600 (80.3%) |
| Character error rate | 16.1% | 5.3% |
| Wrong text fields | 405 | 282 |
| LLM API cost | $0.0012 per form | none |
- Prompt tuning is not the lever. Three rounds on the dev forms landed at 52, 50 and 49 critical fields of 95: noise. The only visible effect was capitalisation.
- The errors are misreads, not inventions: 73% of the wrong fields are one or two characters off, mostly long digit strings. The thin
/of a licence number read as a1in 9 of 10 cases. - The cost, counted fairly: about $0.029 per form of GPU, against about $0.019 for v3e once Chandra’s GPU time is added to its LLM cost. v4a is dearer per form, and it removes the LLM API.
- A second opinion was measured, then rejected. Chandra and Qwen agree on 52.7% of the fields and the agreed ones are 94.4% right, a useful confidence signal. But picking between them gains nothing, and on the digit fields the two agree on only 87 of 240. So the lever was not a second model.
Update: v4b, re-read the hard digit fields from crops
If the prompt cannot fix a misread digit, maybe the model cannot see it. On a whole page a handwritten digit is about half a vision token.
- What changed: the form is a fixed template, so each field sits at a known place. An optional step aligns the page on the template, cuts out six digit fields per vehicle (plate, attestation, policy, licence number, validity dates), enlarges each crop 3 times and sends them in one request to the same server. A value that is read replaces the page value; an unreadable field keeps it.
- Crops checked by eye before the run, and the fields chosen on the 5 dev forms only.
| Metric (same 20 test forms) | v4a | v4b |
|---|---|---|
| Critical fields right | 222 of 380 (58.4%) | 299 of 380 (78.7%) |
| Character error rate | 5.3% | 3.7% |
| Forms with every critical field right | 0 of 20 | 2 of 20 |
| Licence number (of 40) | 12 | 27 |
| Attestation number | 16 | 30 |
| Policy number | 18 | 31 |
| Validity start date | 16 | 34 |
- Pixels were the lever. +77 critical fields with no prompt change and no second model. The thin
/that no prompt could fix became readable. - Of 127 fields changed, 83 became right and 6 became wrong. The step costs about $0.014 per form of GPU.
- First forms ever with every critical field right.
- Write-up: Cropping is all you need in OCR, not prompt tuning covers why prompt tuning failed here and how to apply the idea.
Update: v4c, crop the remaining fields
The same step, over 22 fields instead of 6: names, makes, streets, licence dates, the damage text and the header (date, time, phones).
- The choice was made on dev only: five fields (the insured and driver addresses, the driver first name, the place, the licence prefecture) were worse from crops than from the page, so they were dropped. Crops help on handwriting that is hard to see, not on long free text.
- A run I threw out. The first test pass looked worse than v4b (285 against 299 critical fields). It was not a real result: network errors and a restarting GPU container made the crop call fail on 8 of 20 forms, which fell back to the page read and were saved as if refined. I added retries and made a failed step never saved as done, waited for the server to answer healthy, and reran those forms. The published numbers are from the complete run (20 of 20 forms refined, 0 failures).
| Metric (same 20 test forms) | v3e | v4a | v4b | v4c |
|---|---|---|---|---|
| Critical fields right | 194 (51.1%) | 222 (58.4%) | 299 (78.7%) | 315 of 380 (82.9%) |
| Minor fields right | 381 (63.5%) | 482 (80.3%) | 482 (80.3%) | 501 of 600 (83.5%) |
| Forms with every critical field right | 0 | 0 | 2 | 4 of 20 |
| Character error rate | 16.1% | 5.3% | 3.7% | 2.5% |
| GPU cost per form (estimated) | about $0.019 with the LLM | about $0.029 | about $0.043 | about $0.04 |
- Of 270 fields changed, 156 became right and 37 became wrong. Make went from 37 to 40 of 40, licence expiry from 33 to 38, phone B from 14 to 18.
- Cost, honestly: about 2 times v3e per form, and no LLM API. The GPU figures are wall-clock estimates from one run, and a cold start (reloading 51 GB after 2 idle minutes) can take up to 10 minutes.
What a hard field looks like
Much of what remains is handwriting that is hard to read even for a person (I did not measure a human baseline, so this is an observation, not a number). These are two real crops from a dev form, enlarged 4 times:

| Field | Truth | What the page read | What the crop read |
|---|---|---|---|
| Vehicle B plate | 14472-T-1 | 14472-J-11 | 14472-J-1 |
| Vehicle A licence number | 63/068647 | 631063647 | 63/068647 |
The licence number shows the finding of v4: the thin / was read as a 1 on the whole page and no prompt fixed it, but on an enlarged crop it reads right. The plate shows what is left: the T is written like a J, and cropping does not make that unambiguous.
Where it stands
- Critical fields right: 0.8% to 82.9% (3 to 315 of 380). Minor fields: 2.0% to 83.5%. Character error rate: 89% to 2.5%. Forms with every critical field right: 0 to 4 of 20.
- Marks (ticks, tiles, circles): exact on the synthetic forms, with no model.
- One open 27B model, no fine-tuning, no third-party LLM API. About $0.04 of GPU per form.
- 164 of 980 text fields are still wrong; 16 of 20 forms still have at least one wrong critical field. The weakest fields are the plate, the licence number and the damage text (about 26 to 27 of 40).
- Scored on synthetic forms, which are neater than real ones, and the field positions come from the generator’s template. A person is still needed.
What’s next
- Score all 400 test forms, to replace the 20-form headline with a larger sample.
- Test on real photos, to measure the gap between synthetic and real handwriting and marks.
- Vote or think on the crops, to go after the plate and the licence number.
- A frontier model as the ceiling, to see how far a bigger model goes on the same forms.
- Before production: a throughput test to size the servers, keeping the container warm, and a human-review step for the fields the pipeline is unsure of.
Try it
- Code, tests and CI: github.com/BENHIMA-Mohamed-Amine/constat-ocr
- Every number above, with the run that produced it:
docs/results-log.mdin the repo, one entry per version with the reproduce command.