TeleOCR: A 1.2B Vision-Language Model for Structured Document Parsing

https://huggingface.co/XingChen-AGI/TeleOCR/resolve/main/assets/scientific_figure.png

TeleOCR is an open-source, approximately 1.2-billion-parameter vision-language model for parsing both born-digital documents and camera-captured pages. It is maintained by XingChen-AGI, uses the Transformers library, and is loaded through AutoProcessor and AutoModel. Its central distinction is that it targets geometric distortion and structured content—especially tables and formulas—as well as ordinary text extraction. The README describes deformation-aware document modeling, adaptive sampling, and content-structure decoupled learning; it does not specify the model’s image-resolution limit, training dataset size, training steps, VRAM requirement, or inference speed. Treat the benchmark results as evidence for document-parsing workloads, not as a guarantee for every language, scan quality, or production pipeline.

The project reports strong results on several document benchmarks, including an overall score of 96.87 on OmniDocBench v1.6 and 88.53 on Wild_OmniDocBench. It also reports first place in the ICDAR 2026 Sci-ImageMiner Challenge. These are project-reported evaluations; the supplied material does not provide enough detail to independently...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE