LightOn open-sources LightOnOCR-3 in 0.8B, 1B and 4B sizes under Apache 2.0, scoring up to 86.3 on olmOCR-Bench
Small document-parsing models with a new grounding mode that returns coordinates, suited to turning PDFs and scans into structured text.
// Key points
- Three models at 0.8B, 1B and 4B under Apache 2.0, commercial use allowed; the 0.8B and 4B move to the Qwen3.5 vision-language architecture.
- olmOCR-Bench overall: 86.3 for 4B, 85.5 for 0.8B and 84.5 for 1B, against 87.6 for the 35.1B-parameter Infinity Parser Pro. On ParseBench the 4B scores 75.1, above Infinity Parser Pro's 74.3.
- At 1540 px, throughput peaks at 4.78 pages/s for the 0.8B and 3.36 pages/s for the 4B; the 1B has the lowest single-page latency at 2.7 s.
- A new grounding mode returns labeled bounding boxes and image descriptions, and turns chart data into HTML tables.
Builder's takeThe 0.8B is less than a point behind the 4B, which makes it attractive for document parsing in AI Cloud Drive: I'll run a comparison on the scans and table-heavy PDFs users upload most, to see whether some parsing can move from a cloud API onto our own machines.
// Background · from #Open models
Full timeline →- Oct 9 Liquid AI releases open d1 decision models: d1-3B answers in one forward pass, 8 ms per question on an RTX 4090
- Oct 9 NVIDIA fine-tunes Nemotron to gold-level IOI 2026 (535.4/600) and IMO 2026 (30/42) results, releasing weights and training data
- Oct 7 Mistral Large 4 enters public preview: 1T total, 49B active, $1.36 per million input tokens, weights by month end