LightOnOCR 3 grounds extracted text back to page coordinates
LightOn released LightOnOCR 3 under Apache 2.0 in 0.8B, 1B and 4B sizes. Its grounding mode returns a bounding box for every extracted block, letting a chart value or table cell point back to the exact spot on the scanned page. Grounding adds roughly 25% more tokens, and the pinned vLLM setup breaks the 1B model if Transformers is upgraded manually.
- Three sizes: 0.8B, 1B and 4B, with 4B recommended
- Grounding adds about 25% more tokens than plain transcription
- The 75.1 ParseBench score is LightOn's own metric
- Manually upgrading Transformers breaks the 1B model
Read next
AI