Baidu Open-Sources Unlimited OCR: An Entire Book in One Pass with 32K Context

·Toolin Editorial Team

Baidu has open-sourced Unlimited OCR, which builds on DeepSeek OCR with an R-SWA attention mechanism, letting the model parse dozens of pages in a single forward pass within a 32K standard context — scoring 93.23% end-to-end SOTA on OmniDocBench v1.5.

Baidu Open-Sources Unlimited OCR: An Entire Book in One Pass with 32K Context

DeepSeek OCR pushed visual-token compression to the extreme but left a loose end: with the input side compressed so aggressively, the KV Cache in the decoding stage still balloons as soon as outputs get long, so most OCR systems ultimately fall back to the engineering compromise of "looping over one page at a time." Baidu's newly open-sourced Unlimited OCR picks up exactly where that left off — leaving the encoder untouched and attacking the decoding side, so the model reads an entire book in one forward pass within the 32K standard context length.

The results are solid: on OmniDocBench v1.5, the mainstream document-parsing benchmark, it took end-to-end SOTA with an overall score of 93.23%, a full 6 percentage points above DeepSeek OCR. The model, technical report, and project code are all open-sourced.

Unlimited OCR reads an entire book in one pass

Not page-by-page processing, not for-loop task splitting — dozens of pages of document parsing completed in a single forward pass.

What It Solves

Traditional OCR handles long documents page by page: the model recognizes page one, its state resets, it recognizes page two, and an external scheduler stitches the fragments together. It works as engineering, but semantic coherence is severed — the model has no idea it is performing book-level continuous transcription.

Unlimited OCR's idea comes from an analogy: when a person copies out a book by hand, attention isn't spread evenly across the whole book. Your eyes stay on the source page, your memory holds the few characters you just wrote, and your attention moves on to the next character. You don't fully recall the hundreds of pages already copied while writing the current character.

This "soft forgetting" mechanism is what lets a human copy out an entire book continuously. Unlimited OCR moved it into the attention mechanism.

The Core Innovation: R-SWA Reference Sliding Window Attention

Unlimited OCR is built directly on top of DeepSeek OCR, with 3B total parameters and 500M activated parameters, reusing DeepEncoder for aggressive visual-token compression (a 1024×1024 page reduces to just 256 visual tokens). Baidu's key change is replacing standard multi-head attention (MHA) with R-SWA (Reference Sliding Window Attention).

R-SWA splits what the model can see into two parts:

  • Reference tokens: the source document's visual tokens + prompt, kept permanently "in view" with no fidelity decay over time
  • The most recent 128 output tokens: only the short stretch just generated is kept as working memory, mimicking "remembering only the last few characters written"

R-SWA: every generated token attends to all reference tokens (visual + prompt) and the previous n=128 output tokens. KV cache stays constant throughout the entire decoding process.

Why This Is the Key

An ordinary sliding window slides the visual tokens out too, so visual features gradually blur over long generations; standard MHA, meanwhile, lets the KV cache grow with output length, getting slower and more memory-hungry the further you go.

R-SWA locks the visual tokens in the reference cache and confines output tokens to a fixed window, so:

  • KV cache stays constant throughout — no more linear growth with output length
  • Decoding latency stays constant — per-call time on Flash Attention v3 kernels is essentially a flat line, while the DeepSeek OCR baseline shows clear latency spikes
  • GPU memory usage is fixed

Benchmark Results

Main Benchmark: OmniDocBench v1.5

Unlimited OCR took end-to-end SOTA at 93.23%. It didn't fall behind on complex-layout documents like PPTs, newspapers, magazines, and notes either, showing that R-SWA isn't just friendly to plain text.

No drop-off on complex layouts either.

Long-Document Parsing: 40+ Pages in One Go

Baidu built an internal long-document test set grouped by page count — 2/5/10/15/20/40+ pages:

  • Inputting 20 pages at once still holds up well
  • In the 40+ page scenario, edit distance stays below 0.11 and Distinct-35 is about 97% (not prone to repetitive output)

Long-Output Speed Advantage

At 256 output tokens, the two models are nearly identical in speed; by 6000 tokens, DeepSeek OCR's TPS has fallen 35% behind Unlimited OCR. The longer the output, the more obvious R-SWA's advantage.

MHA slows down as the KV cache grows; R-SWA locks the output-side KV cache to a fixed window, so decoding overhead doesn't balloon with output length.

How to Use It

The open-source resources are complete:

  • Technical report: Unlimited-OCR Works (readable as a PDF on HuggingFace)
  • GitHub: github.com/baidu/Unlimited-OCR
  • Hugging Face: huggingface.co/baidu/Unlimited-OCR

The model is based on DeepSeek OCR — think of it as "swapping the decoder-side attention on DeepSeek OCR." If your workflow already uses DeepSeek OCR, the migration cost is mostly replacing the attention layer.

Use Cases

  • Whole-book / long-report digitization: transcribe an entire PDF in one forward pass, no page-by-page scheduling
  • Batch parsing of academic material: cross-page formulas, figures, and references keep their semantic continuity
  • Archives / newspapers / magazines: complex multi-column layouts finished within a single inference
  • RAG document preprocessing: long-document parsing no longer gets its semantics truncated by an external stitcher

💡 Tip: The HuggingFace acknowledgments explicitly thank "DeepSeek-OCR, DeepSeek-OCR-2," and the technical report cites DeepSeek OCR as many as 40 times. If your team is already in the DeepSeek OCR ecosystem, Unlimited OCR is an upgrade along the same road, not a stack change.

Project page: github.com/baidu/Unlimited-OCR

Related articles

OpenWorker: Andrew Ng's Open-Source Desktop AI Coworker That Ships Finished Work
AI Products

OpenWorker: Andrew Ng's Open-Source Desktop AI Coworker That Ships Finished Work

Andrew Ng open-sources OpenWorker (MIT), a local-first, model-agnostic desktop agent connecting to 25+ tools, turning "prepare a customer brief" into a finished document you can actually open.

Toolin Editorial Team
AutoMIA: Two Images In, a 3D-Printable Mirror Illusion Out — Runs on a Single RTX 3090
AI Products

AutoMIA: Two Images In, a 3D-Printable Mirror Illusion Out — Runs on a Single RTX 3090

AutoMIA, a Tsinghua CVPR'26 Highlight open-source project, turns two input images into a 3D-printable voxel mirror-illusion art model (STL-ready) — single RTX 3090, 76 seconds per design.

Toolin Editorial Team
Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL
AI Products

Macaron-V1: An Open-Source Coding Model That Compresses Personalized Experience into LoRAs with MoL

Mind Lab open-sources Macaron-V1 (Venti / Coding-Venti / Tall tiers), a Mixture-of-LoRA architecture built on GLM-5.2 744B plus four 1B LoRA experts, with a free hosted API and a UI4A plugin.

Toolin Editorial Team
UniWorld-View: Single Image/Video to Any Camera Trajectory, Tops the WorldScore Leaderboard, Fully Open-Sourced
AI Products

UniWorld-View: Single Image/Video to Any Camera Trajectory, Tops the WorldScore Leaderboard, Fully Open-Sourced

UniWorld-View, open-sourced by Peking University + Rabbitpre + Pengcheng Laboratory, generates novel-view videos along any camera trajectory from a single image/video — it tops Fei-Fei Li's team's WorldScore leaderboard, with code and weights fully open under Apache-2.0 and one-click downloads.

Toolin Editorial Team
Claude Managed Agents Ships Six Updates: Skill Cap Raised to 500, With Ready-to-Run Payloads
AI Products

Claude Managed Agents Ships Six Updates: Skill Cap Raised to 500, With Ready-to-Run Payloads

Anthropic's managed agent platform CMA ships six updates at once: per-session skills up from 20 to 500, a five-level effort setting writable into per-agent config, and seeded sessions created in one step with initial_events.

Toolin Editorial Team
DojoAgents: Build a Financial Research Agent Locally in 10 Minutes
AI Tutorials

DojoAgents: Build a Financial Research Agent Locally in 10 Minutes

An open-source agent framework from Shenchong Intelligence covering A-shares, US stocks, and Hong Kong stocks — deploy a financial research agent that autonomously analyzes market themes, locally, in 10 minutes.

Toolin Editorial Team