Baidu Open-Sources Unlimited OCR: a 500M-Active Small Model Reads 40 Pages in One Pass Without Forgetting

·Toolin Editorial Team

Baidu open-sources Unlimited OCR, an end-to-end OCR model with 3B total / 500M active parameters that sets a new OmniDocBench SOTA and transcribes dozens of pages per inference without forgetting.

Baidu Open-Sources Unlimited OCR: a 500M-Active Small Model Reads 40 Pages in One Pass Without Forgetting

If you've done document digitization, contract parsing, or long-PDF transcription, you've probably been tormented by one shared problem: the OCR model "forgets" page by page, and a dozens-of-pages document gets stitched together by an external scheduler, getting slower and messier the further it goes. Baidu's newly open-sourced Unlimited OCR targets this pain point directly — a single inference reads from the first page to the last, with constant KV-cache occupancy.

It suits developers who need end-to-end long-document parsing, teams building RAG knowledge bases, and anyone who wants OCR inside a production pipeline without being dragged down by "page-by-page processing." The model and code are fully open-sourced and ready to use right now.

Unlimited OCR: an end-to-end OCR model with 3B parameters / 500M active

What Unlimited OCR Is

One-line summary: an end-to-end OCR MoE model that takes "reference sliding window attention" (R-SWA) to its extreme.

  • 3B total parameters, only 500M actually active — nearly negligible in the era of large models
  • Scores 93.23% on OmniDocBench v1.5 and 93.92% on v1.6, a new end-to-end SOTA
  • Among the rivals on the same stage, the 235B Qwen3-VL scores 89.15, the 72B Qwen2.5-VL scores 87.02, and Gemini-2.5 Pro scores 88.03

With fewer active parameters than their pocket change, it leaves them all behind on the benchmarks. That is Unlimited OCR's most obvious selling point.

Core Mechanism: R-SWA Solves "Page-by-Page Amnesia"

Under standard attention, the KV cache snowballs as output grows — memory can't take it and speed keeps dropping. That is the real reason all OCR models are forced into page-by-page processing and frequent "amnesia."

Baidu's answer is Reference Sliding Window Attention (R-SWA), which mirrors how a person copies from a book:

  • For every token generated, the model looks at all the "reference tokens" (the visual tokens of the whole image + the prompt), so it always "sees" the complete original
  • But on the output side it only looks back at the previous 128 tokens — like glancing at the last few lines you just wrote while copying a book
  • Once all attention layers switch to R-SWA, the KV cache becomes a fixed-capacity queue, and outputting 10,000 tokens costs exactly the same memory as outputting 100,000

R-SWA keeps the KV cache constant, with output latency a flat line throughout

Paired with the DeepEncoder that first appeared in DeepSeek OCR, a 1024×1024 PDF page is compressed to just 256 visual tokens (a 16x compression ratio). Because visual tokens don't participate in state transitions under R-SWA, the image information stays crystal clear no matter how long the document, never degrading through decoding.

Dozens of Pages Read in One Inference

Within a standard 32K context, Unlimited OCR transcribes dozens of pages in a single forward pass:

Input documentEdit distance (word-by-word against the original)Distinct-35 repetition
20 pages0.057—
40+ pages< 0.1197%

Dozens of pages transcribed in one breath, with almost no parroting.

On benchmark comparisons, OmniDocBench v1.5 overall score of 93.23% beats DeepSeek OCR's 87.01% by 6.22 percentage points; text edit distance drops from 0.073 to 0.038, formula CDM leaps from 83.37 to 92.61, and table TEDS rises from 84.97 to 90.93. Across the nine document types (PPT, academic papers, magazines, newspapers, etc.), it surpasses DeepSeek OCR on both text and reading order across the board, and leads DeepSeek OCR 2 in seven categories.

Efficiency comparison: Unlimited OCR's TPS stays above DeepSeek OCR's throughout

Efficiency is a rout too: at 6144 output tokens, Unlimited OCR's TPS is 7847 while DeepSeek OCR has fallen to 5822 — a 35% gap. And this is a 500M-active MoE small model trained just 4000 more steps on top of DeepSeek OCR — R-SWA is nearly a "free lunch" for parsing tasks.

How to Use It

The repository and model weights are both open:

Clone the repo and deploy following the README; the weights support local inference. For teams that want OCR inside a production pipeline without the burden of page-by-page scheduling logic, this is a new path worth evaluating.

Where It Fits

  • End-to-end long-document transcription: contracts, papers, reports read in one pass, with no external page-splitting scheduler
  • RAG knowledge base construction: stable full-text output makes chunking and vectorization easier
  • Tables / formulas / mixed-layout parsing: clearly ahead on the TEDS and CDM metrics
  • Edge / low-compute deployment: a 500M-active footprint is hardware-friendly

💡 Tip: The current version has a 32K context window; the paper's outlook mentions training to 128K next and building a prefill pool so the model learns to turn pages automatically — at which point OCR's boundary would extend from "recognizing one page of text" to "understanding an entire book."

The People Behind It

An interesting detail in the technical report's author list: three core contributors, with the technical director credited only under the initials "YY." In the GitHub acknowledgments, Deepseek-OCR and Deepseek-OCR-2 are listed first and second; combining the capabilities, the timeline, and the signature style, outside observers widely speculate that YY is Wei Haoran, the OCR core author who left DeepSeek (he single-handedly built DeepSeek OCR from generation one to two, including DeepEncoder and the MoE decoder).

For users, the more practical takeaway: if this R-SWA paradigm spreads further to ASR (speech recognition) and translation, what Baidu holds won't just be an OCR model but a general-purpose long-form parsing framework. Worth watching over the long haul.