olmOCR-Bench

Overall score versus model parameters. Smaller and higher is better.

olmOCR-Bench score versus model sizeLower parameter count and higher overall score are better. Dashed line: Pareto frontier.75808590950.712481632Model parameters (B) · Smaller is betterOverall score · Higher is betterMistral OCR 4 — 85.20 points — Size undisclosed — Reported resultsMistral OCR 4† · 85.2Infinity-Parser2-Pro — 87.6 points — 35.1B — Reported resultsInfinity-Parser2-Pro†LightOnOCR-3-4B — 86.3 points — 4.0B — Output with post-processingLightOnOCR-3 4B†chandra-ocr-2 — 85.8 points — 4.0B — Output with post-processingchandra-ocr-2†LightOnOCR-3-0.8B — 85.5 points — 0.8B — Output with post-processingLightOnOCR-3 0.8B†LightOnOCR-3-1B — 84.5 points — 1.0B — Output with post-processingLightOnOCR-3-1B†dots.mocr — 83.9 points — 3.0B — Reported resultsdots.mocrjina-ocr-v1 — 83.4 points — 3.4B — Reported resultsjina-ocr-v1surya-ocr-2 — 83.3 points — 0.7B — Output with post-processingsurya-ocr-2†LightOnOCR-2-1B — 83.2 points — 1.0B — Reported resultsLightOnOCR-2 *
External modelsLightOnOCR-2LightOnOCR-3Pareto frontier† Processed / reported olmOCR results
Hover or focus a point for its exact checkpoint and scores.
Method and source

Smaller size and higher score are better. The frontier includes points that no other comparable point matches or improves on in both dimensions, with a strict improvement in at least one. Comparisons use source precision; displayed scores are rounded to one decimal. Parameter counts are those in the tables, including nominal backbone sizes where specified.

Overall uses five categories for ParseBench and includes headers/footers for olmOCR. All points match the selected table rows; LightOnOCR-3 uses the candidate-1 evaluations for 0.8B, 1B, and 4B. Postprocessing applies only to olmOCR; ParseBench and fr-bench-pdf2md outputs are raw. The same candidate checkpoints are used across all three benchmarks. Missing scores or sizes are omitted. Models with undisclosed parameter counts (such as MistralOCR4.1) cannot be placed on a size axis. * LightOnOCR-2 uses its published olmOCR score excluding headers/footers; it is shown but excluded from the frontier. Mistral OCR 4 is a horizontal score reference because its size is undisclosed; it is excluded from the size frontier. Its olmOCR score is the user-provided reported 85.20; ParseBench and fr-bench-pdf2md retain their existing selected scores.

Leaderboard source revision: 3576c5240602588f596f8790a3a7161a524e59ad. Values are pinned to this snapshot.

Show scores and output details
Plotted scores
ModelSize (B)OverallOutput / inputsFrontier
Infinity-Parser2-Pro35.187.6Reported results✓
LightOnOCR-3-4B4.086.3Output with post-processing✓
chandra-ocr-24.085.8Output with post-processing—
LightOnOCR-3-0.8B0.885.5Output with post-processing✓
Mistral OCR 4Undisclosed85.2Reported resultsSize undisclosed
LightOnOCR-3-1B1.084.5Output with post-processing—
dots.mocr3.083.9Reported results—
jina-ocr-v13.483.4Reported results—
surya-ocr-20.783.3Output with post-processing✓
LightOnOCR-2-1B1.083.2 *Reported resultsExcluded *