TeleOCR: a 1.2B vision-language model for structured document parsing
XingChen-AGI released TeleOCR, an open-source ~1.2B-parameter vision-language model for parsing born-digital and camera-captured documents. It scored 96.87 on OmniDocBench v1.6 and 88.53 on Wild_OmniDocBench, and took first place in the ICDAR 2026 Sci-ImageMiner Challenge. It is licensed under Apache-2.0.
- ~1.2B parameters versus 30B–241B general VLMs
- 96.87 on OmniDocBench v1.6 and 88.53 on Wild_OmniDocBench
- Table TEDS 97.05 and Table TEDS-S 98.52
- Parses distorted DocUNet and DIR300 pages without dewarping
Read next
AI