TeleOCR Parses Crumpled, Photographed Pages Without Dewarping

A research team at China Telecom’s AI subsidiary (TeleAI) has released TeleOCR, a 1.2-billion-parameter vision-language model that parses both clean digital documents and photographed ones — folded, curled, or crumpled — with no separate dewarping step. The weights first went up on Hugging Face on August 17, 2026 under the name NaviDC-OCR. The team renamed the project TeleOCR on September 10, 2026, alongside version 3 of its technical report. The model is released under Apache 2.0, and the team reports an overall score of 96.87 on OmniDocBench v1.6, the highest in its own comparison table.

Advanced

A grid of twelve photographs of folded, curled and crumpled paper documents, each overlaid with coloured polygons marking text blocks, titles and figures detected by TeleOCR
TeleOCR layout predictions on photographed pages from the DIR300 dewarping dataset, run without rectification. Image credit: TeleOCR on GitHub

Two Kinds of Document, One Model

Document parsing models have mostly been tuned for one of two settings. Digital PDFs and clean scans have straight lines and rectangular layout regions. Photos taken with a phone are bent, skewed, and unevenly lit. According to the paper, both main families of parser struggle with the second setting, for different reasons. Decoupled parsers detect layout first and then read each region, so a bent page throws off the layout stage and “can introduce cascading errors”. End-to-end parsers skip explicit layout detection but “often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios.”

TeleOCR stays decoupled, with layout first and recognition second, but it teaches the model to handle distortion directly. It does not rely on a rectification model to flatten the page first.

Technical Details

The model combines a vision encoder taken from Qwen2.5-VL, a Qwen3-0.6B language model, and an MLP aligner trained from scratch, for about 1.2B parameters in total. That puts it in the same weight class as MinerU2.5 (1.2B) and PaddleOCR-VL (0.9B). A single set of weights handles several tasks, each selected by a text prompt: plain text, tables (emitted in the compact OTSL table markup), LaTeX formulas, code blocks, page layout, polygon layout for distorted pages, and recovering the underlying data table from a scientific chart.

Diagram of TeleOCR's four training stages: document parsing pre-training, deformation-aware training with region-level and point-level awareness, content-structure decoupled learning for tables and formulas, and reinforcement learning
TeleOCR’s four-stage training pipeline. Image credit: TeleOCR on GitHub

Training runs in four stages:

  • Pre-training on captioning, interleaved image-text, and OCR data.
  • Deformation-aware training on 4M digital samples plus 2M synthetically distorted ones. For bent pages, the model predicts polygon outlines for each layout region instead of rectangular boxes. A second, point-level task has it predict an M×M grid of control points showing how the page surface is warped.
  • Content-structure decoupled learning on about 120K curated samples. The model first learns table and formula structure separately from their content, for example OTSL cell tokens with the text stripped out, or a LaTeX skeleton with the symbols removed.
  • Reinforcement learning with GRPO and task-specific rewards.

Two data-side techniques support this. Curvature-Guided Douglas–Peucker Sampling (CGDP) decides where to place polygon vertices: it weights each candidate point by local curvature, so sharp creases get more vertices than gentle bends. Multi-node Consensus Voting produces pseudo-labels by running several different parsers on the same unlabelled page. It keeps the output that agrees most with the others and discards pages where agreement falls below a threshold.

Bar charts comparing TeleOCR with other document parsers on OmniDocBench v1.6, Wild_OmniDocBench and PureDocBench, with TeleOCR highest on each
Benchmark comparison as reported by the TeleOCR team. Image credit: TeleOCR on GitHub

The Numbers

All figures below come from the TeleOCR model card and paper:

  • OmniDocBench v1.6 (digital documents): 96.87 overall, ahead of OvisOCR2 (96.58), PaddleOCR-VL-1.6 (96.33), MinerU2.5-Pro (95.75) and GLM-OCR (95.22). Table TEDS is 97.05, against 94.76 for the next-best model. TeleOCR does not lead on every sub-metric: OvisOCR2 has lower text edit distance (0.025 vs 0.027) and higher formula CDM (97.53 vs 96.36).
  • Wild_OmniDocBench (photographed documents): 88.53 overall, against 87.91 for OvisOCR2 and 87.36 for PaddleOCR-VL-1.6.
  • PureDocBench: an average of 78.41 across clean, digitally degraded, and real-degraded pages. On the real-degraded subset alone, Gemini-3.1-Pro scores slightly higher (71.98 vs 70.85).
  • ICDAR 2026 Sci-ImageMiner: first place, with a weighted score of 41.81.

The team also reports 67.96 on the EMNLP 2026 Dr.DocBench Challenge, compared with 62.26 for MinerU 2.5 Pro. The model card says it ran this evaluation itself, “with its native weights”. That makes it a self-run comparison, not a leaderboard placement. On the digital benchmark, the top four models are within about half a point of each other, and all of them are under 1.3B parameters.

What This Means

The main contribution is the camera-captured side rather than the OmniDocBench lead. Scores on clean PDFs are now tightly bunched among sub-1.3B models, and less than a third of a point on that benchmark separates TeleOCR from the runner-up. Phone photos of handouts, receipts, and bound books are the harder case, and the usual workaround has been a separate dewarping model before OCR. TeleOCR’s polygon-layout approach removes that stage. For anyone digitising physical archives or processing photographed paperwork, one fewer model in the pipeline is a practical gain.

Some caution is warranted. Every comparison in the release is the authors’ own. In the OmniDocBench v1.6 table on the model card, the sub-metrics listed for HunyuanOCR-1.5 are identical to PaddleOCR-VL-1.6’s, which looks like a transcription error. Teams choosing a parser should confirm competitor numbers against the benchmark’s own repository. The release has also appeared under three names and several Hugging Face organisations (StarDoc-AI, XingChen-AGI) in six weeks, and the Quick Start code on the model card still loads from StarDoc-AI/TeleOCR. Practically, the model is small enough for a single consumer GPU. The team documents a vLLM-based pipeline for full-document parsing, and a community GGUF conversion adds llama.cpp support.

Related Coverage

This post was drafted with AI assistance and reviewed by RITS staff.

Sources