Baidu Open-Sources Unlimited OCR: Reading an Entire Book in One Forward Pass

·Toolin Editorial Team

Unlimited OCR, Baidu's model built on DeepSeek OCR, achieves 32K-context long-document parsing through the R-SWA mechanism, with an end-to-end SOTA of 93.23% on OmniDocBench v1.5.

Baidu Open-Sources Unlimited OCR: Reading an Entire Book in One Forward Pass

DeepSeek OCR left behind a hard problem in long-document parsing, and Baidu has picked it up. On HuggingFace, Baidu open-sourced a new model, Unlimited OCR, which under a standard maximum context length of 32K lets an OCR model read an entire book in one go for the first time — not page-by-page processing, not for-loop-style task splitting, not stitching results together with an external scheduler, but genuinely parsing dozens of pages of a document in a single forward pass. This piece breaks down what problem it solves, how to use it, and what can be done with it now that it is open-sourced.

What Unlimited OCR Is

Think of it as "DeepSeek OCR Plus that reads long documents." It is built directly on top of DeepSeek OCR — the visual compression side was already pushed to the extreme by DeepSeek OCR (a 1024×1024 document page encodes down to just 256 visual tokens), so Unlimited OCR did not rebuild the encoder; it put all its effort into the decoding stage. The project homepage states the thesis in one line: "push DeepSeek-OCR one step further".

The Core Problem It Solves

The Old Models' Bottleneck: KV Cache Bloat on the Decode Side

Why is it still so hard for DeepSeek OCR to handle long documents even with such aggressive visual token compression? The answer is on the decode side.

After visual token compression, the text the model generates does not vanish into thin air. As the output grows longer, the KV cache in the decoder keeps growing: the longer the output, the higher the VRAM usage; the longer the history, the heavier the attention computation; generation gets slower and slower.

This is why most OCR systems in the past ultimately fell back to page-by-page parsing — no matter how efficient the encoder, it cannot solve the ever-growing historical burden in the decode stage.

Unlimited OCR parses dozens of document pages in a single forward pass

The key is not "can it handle multiple pages" but "can it do so without degrading into page-by-page mode" — Unlimited OCR chose the latter.

The R-SWA Mechanism: Controlling Cost Without Losing Long-Range Dependencies

The core mechanism Unlimited OCR introduces is R-SWA (Rotary Sliding Window Attention). The problem it solves is blunt: control the attention compute cost without sacrificing long-range dependency modeling. Simply put, the model can still see content from dozens of pages back without compute exploding exponentially with page count — this is the underlying reason it can run through an entire book in one pass.

Core Capabilities and Benchmark Performance

Performance Numbers

On OmniDocBench v1.5, the mainstream document-parsing benchmark:

  • Unlimited OCR takes end-to-end SOTA with a total score of 93.23%
  • A full 6 percentage points above DeepSeek OCR

OmniDocBench v1.5 is the standard benchmark in the document-parsing field; 93.23% means state-of-the-art end-to-end extraction on real documents (papers, textbooks, contracts, and so on).

Where It Fits

Because a single forward pass covers dozens of pages, Unlimited OCR is especially suited to:

  • One-shot parsing of whole books / long reports: this used to mean page-by-page runs followed by manual stitching; now it is one pass
  • Cross-page consistency: charts, references, and glossaries cited across pages are no longer interrupted by page-by-page processing
  • Structured document extraction: complex tables, formulas, and layouts in long contracts and long papers can be read whole before extraction

How to Use It

The model is open-sourced on HuggingFace; the weights and code can be pulled directly:

  • Model page: search "Unlimited OCR" on HuggingFace (official Baidu release)
  • Base: built on DeepSeek OCR, with good compatibility with the original project
  • Deployment: loads through the standard transformers / vLLM flow, no special frameworks needed

💡 Tip: Because it is built on top of DeepSeek OCR, engineering teams already using DeepSeek OCR can migrate smoothly — mainly swapping the attention implementation and loading the new weights, with no need to redo the data pipeline.

Hands-On Impressions

Strengths

  • Truly long-range: parses an entire book in one forward pass, ending the engineering complexity of page-by-page stitching
  • Solid foundation: standing on the shoulders of DeepSeek OCR's already-extreme compression, engineering migration cost is low
  • Fully open: weights and code are both open — deploy locally and fine-tune further

Limits

  • Hardware threshold: 32K-context long-range inference still demands VRAM; local deployment needs a GPU of the appropriate spec
  • Domain fit: SOTA on benchmarks does not mean optimal for every vertical document type (say, highly specialized medical imaging); self-test on critical scenarios

Use Cases

  • Publishing and academia: structured extraction of entire textbooks and entire papers — building knowledge bases without page-by-page processing
  • Legal and compliance: end-to-end clause extraction from long contracts and long case files
  • Enterprise document platforms: replace traditional page-by-page OCR pipelines and cut engineering complexity
  • Personal long-document processing: deploy locally and convert an entire PDF to structured text in one pass