HunyuanOCR-1.5 Hands-On: SOTA End-to-End OCR from a 1B-Parameter Model

·Toolin Editorial Team

Tencent Hunyuan's HunyuanOCR-1.5 packs document parsing, text recognition, information extraction, and image-text translation into a single 1B-parameter VLM, and pushes inference speed up 6x. This guide walks you through running it locally or on vLLM.

HunyuanOCR-1.5 Hands-On: SOTA End-to-End OCR from a 1B-Parameter Model

If you work on OCR tasks like invoice recognition, table extraction, scanned ancient texts, or contract parsing, you have probably suffered through the traditional OCR pipeline's four-stage grind of "detection → recognition → layout → structuring." HunyuanOCR-1.5, open-sourced by Tencent Hunyuan on 2026-07-07 (also called HyOCR-1.5 in the repo), offers a cleaner path: it packs all of the above into a single 1B-parameter vision-language model (VLM) that produces end-to-end results in one pass, and uses speculative decoding to speed up long-document decoding by roughly 6.37x.

This article walks you through getting the official repository running, from downloading the weights to a batch inference run that outputs Markdown, and covers the two most common paths: a vLLM server and a local llama.cpp deployment. All commands come from the official Tencent-Hunyuan/HunyuanOCR repository — nothing here is made up.

Before You Start

💡 Note: the official docs explicitly state that the three inference environments — vLLM (AR / DFlash) and native transformers — are mutually incompatible due to transformers version conflicts, and must be set up separately. Pick one of paths A/B/C below and don't mix them.

Step by Step

Step 1: Download the model weights

The weights are hosted on Hugging Face at tencent/HunyuanOCR; the download includes both the base model and the DFlash draft model.

pip install -U "huggingface_hub[cli]"
huggingface-cli download tencent/HunyuanOCR \
    --local-dir ./HunyuanOCR --exclude "v1.0/*"

The --exclude "v1.0/*" flag keeps you from also pulling down the previous-generation 1.0 weights, saving disk space.

Step 2: Pick an inference path

The official repo provides three mutually independent environments under inference/, plus a PC-side llama.cpp option:

PathFrameworkDFlash accelerationBest for
A. inference/vllm_0_18_1vLLM 0.18.1NoSimplest; pure AR inference
B. inference/nightlyvLLM nightly (CUDA 13)YesLong documents/tables/formulas, when you need every bit of speed
C. inference/transformersHF transformers 5.13—Alignment/accuracy checks
D. llama.cppGGUF + OpenAI-compatible serverOptionalCPU / laptop / consumer GPU

Below we run path A (the most common single-GPU vLLM setup) end to end; if you only care about the speedup, just follow path B's README and swap serve.sh for serve_dflash.sh.

Step 3: Start the vLLM server

First install the dependencies from inference/vllm_0_18_1/requirements.txt, then bring up an OpenAI-compatible server:

MODEL_PATH=./HunyuanOCR GPU=0 PORT=8000 \
    bash inference/vllm_0_18_1/serve.sh

# Health check
curl -sf http://127.0.0.1:8000/v1/models

The server exposes the model as tencent/HunyuanOCR, with --max-model-len defaulting to 131072 (128K context, matching the maximum configuration used in training).

Step 4: Document parsing on a single image

The official tooling locks task types down to 12 --task-type values (use --list-tasks to see them all), so an off-the-cuff prompt can't steer the model astray. Sampling parameters (temperature=0.0, top_p=1.0, repetition_penalty=1.08) and early stopping on trailing repetition are already built into the client.

python inference/vllm_0_18_1/infer_vllm_client.py \
    --image /path/to/document.png \
    --task-type doc_parse \
    --model tencent/HunyuanOCR \
    --port 8000 --max-tokens 32768

The most commonly used task types: doc_parse (parse a full page into structured text), text_spotting (text-only localization and recognition), information_extraction (extract fields per instructions), and text_image_translation (image-text translation). One image in, one block of Markdown or JSON out, depending on which task you pick.

Step 5: Batch inference

When you're facing a pile of scanned documents, use the official batch_infer.py, which ships with multi-endpoint concurrency and resumable runs:

python inference/vllm_0_18_1/batch_infer.py \
    --image-dir /path/to/images \
    --out-dir /path/to/output \
    --ports 8000 \
    --task-type doc_parse \
    --max-tokens 32768 \
    --concurrency 16

If you have multiple GPUs, start several serve.sh instances and pass the port list to --ports for linear throughput scaling.

Running on a Laptop: the llama.cpp Path

It works without a GPU too. HunyuanOCR-1.5 provides GGUF-converted checkpoints and an OpenAI-compatible llama-server; community llama.cpp supports the base model, while DFlash acceleration requires an adapted fork (wendadawen/llama.cpp @ dflash-adapt-hunyuanocr-hunyuanstyle).

Minimal workflow (community build, no DFlash):

# 1. Build llama.cpp
git clone https://github.com/ggml-org/llama.cpp.git && cd llama.cpp
cmake -B build -DLLAMA_BUILD_EXAMPLES=ON   # add -DGGML_CUDA=ON for NVIDIA GPUs
cmake --build ./build --config Release -j

# 2. Convert to GGUF (base + mmproj)
python3 convert_hf_to_gguf.py --outfile ./HunyuanOCR/hyocr-f16.gguf        --outtype f16 ./HunyuanOCR
python3 convert_hf_to_gguf.py --outfile ./HunyuanOCR/mmproj-hyocr-f16.gguf --outtype f16 --mmproj ./HunyuanOCR

# 3. Start the OpenAI-compatible server
build/bin/llama-server \
    --model  ./HunyuanOCR/hyocr-f16.gguf \
    --mmproj ./HunyuanOCR/mmproj-hyocr-f16.gguf \
    --host 0.0.0.0 --port 8080 --alias HYVL \
    --ctx-size 10240 --n-predict 4096

The full DFlash-adapted build, draft weight conversion, and smoke-test scripts for the 26 sample images are all in the official docs/llama_cpp.md.

Why Version 1.5 Is Worth Switching To

Two core changes explain the "faster and better":

  • DFlash speculative decoding: long autoregressive decoding is the biggest bottleneck for end-to-end OCR (dense documents, tables, and formulas easily run to thousands of tokens). HunyuanOCR-1.5 uses a lightweight block-diffusion draft model to draft multiple candidate tokens in parallel, which the target model then verifies in one shot. Decoding latency for long structured outputs drops significantly while leaving the target model's output distribution unchanged (lossless acceleration). Official speed claims on OCR tasks are in docs/benchmark.md.
  • Agentic Data Flow + an upgraded training recipe: on the data side, an agent-driven data construction system translates the model's weak spots into actionable data requirements, specifically reinforcing long-tail capabilities such as low-resource OCR, ancient-text OCR, and multi-image text QA; on the training side, Stage-3 was replanned, image resolution pushed to 4K, the context window extended to 128K, and RL applied across different OCR tasks in post-training.

Verifying the Results

After finishing step 4, you should get a block of structured Markdown / JSON on stdout: tables keep their rows and columns, formulas come out as LaTeX, and multi-column layouts are restored in reading order. To compare accuracy further, you can use the two open-source benchmarks released by the team alongside: Chronicles-OCR (perception of Chinese "seven scripts" ancient writing, arXiv:2605.11960) and ChartArena (cross-language chart parsing, arXiv:2606.01348).

FAQ

  • Can the three inference environments share one conda env? No. The transformers versions required by vLLM and native transformers are mutually incompatible; the official inference/README.md stresses that this is "a verified limitation, not a preference." Set up a separate env per path.
  • vLLM results don't exactly match transformers? The README confirms that the vLLM framework has had accuracy differences from transformers since the 1.0 era, and the team is still fixing them. For accuracy alignment, treat inference/transformers as the source of truth.
  • Is commercial use free? The repository license is the Tencent Hunyuan Community License Agreement, friendly to personal and research use; for commercial use, read the full LICENSE first.
  • Can it be fine-tuned? Yes. scripts/sft_base.sh is full-parameter SFT, scripts/sft_dflash.sh trains a DFlash draft from scratch, and scripts/sft_dflash_finetune.sh continues fine-tuning an existing draft; see docs/training.md for the exact hyperparameters.

Related articles

Sakana Fugu: the Orchestrator That Doesn't Answer Itself, Just Directs Other Models to Do the Work
AI Products

Sakana Fugu: the Orchestrator That Doesn't Answer Itself, Just Directs Other Models to Do the Work

Sakana AI releases the Fugu family of orchestrator models, smartly dispatching GPT, Claude, and Gemini to finish tasks, with performance approaching Fable 5 and Mythos Preview.

Toolin Editorial Team
Seko Infinite Canvas Hands-On: Drop in One Idea, and the Agent Finishes a Whole AI Video for You
AI Tutorials

Seko Infinite Canvas Hands-On: Drop in One Idea, and the Agent Finishes a Whole AI Video for You

A hands-on guide to the Seko infinite canvas + Seedance 2.0: 720P cost down 50% and 1080P down 80%, letting even beginners produce multi-episode AI video epics in 10 minutes.

Toolin Editorial Team
WeChat Xiaowei Beta Hands-On: Swipe Right and Turn All of WeChat into an Agent
AI Products

WeChat Xiaowei Beta Hands-On: Swipe Right and Turn All of WeChat into an Agent

A beta look at WeChat's official AI assistant Xiaowei: chat summaries, auto-replies, invoking mini programs, money transfers, Moments browsing, and building small tools — eight capabilities in one read.

Toolin Editorial Team
The Ministry of Education's "Sunshine Volunteer" AI Assistant: Free College-Application Plans Generated in Seconds
AI Products

The Ministry of Education's "Sunshine Volunteer" AI Assistant: Free College-Application Plans Generated in Seconds

The Ministry of Education has officially upgraded its "Sunshine Volunteer" system, with the "Zhihui Xiaozhao" AI assistant online 24 hours a day, providing reach-match-safety application plans free of charge based on official data.

Toolin Editorial Team
GenShield: An All-in-One Open Source Framework for AI-Generated Image Detection and Repair
AI Products

GenShield: An All-in-One Open Source Framework for AI-Generated Image Detection and Repair

A Peking University team has open sourced GenShield, unifying AI-generated image detection and artifact correction in a single autoregressive framework with 98.8% detection accuracy

Toolin Editorial Team
OpenAI Codex OSS Mode: Connect Local Models with a Single Line of Config
AI Products

OpenAI Codex OSS Mode: Connect Local Models with a Single Line of Config

Codex adds an OSS mode supporting local model services like Ollama and LM Studio, enabling offline operation and cost control

Toolin Editorial Team