FineVLA: Which Hand, Where to Grasp — One Sentence Sets It All
HKU's XLANG Lab and Alibaba's Qwen team have open-sourced FineVLA, a framework that refines language supervision in robot policy learning from "what to do" to "how to do it," reaching 86.8% success in RoboTwin simulation (+15) and 62.7/100 on a real dual-arm robot, with code, models, and benchmark all open-sourced.


FineVLA: Which Hand, Where to Grasp — One Sentence Sets It All
HKU's XLANG Lab and Alibaba's Qwen team have open-sourced FineVLA, a framework that refines language supervision in robot policy learning from "what to do" to "how to do it," reaching 86.8% success in RoboTwin simulation (+15) and 62.7/100 on a real dual-arm robot, with code, models, and benchmark all open-sourced.
Tell a robot to "put the cup in the basket" and it can complete the task — but which hand to use, which direction to approach from, whether to grasp the cup's body or its handle: these execution-defining details are rarely labeled in existing robot datasets. So what the model learns is "the end result must succeed," while the execution constraints remain hard to learn from language. FineVLA, a framework jointly proposed by the XLANG Lab at the University of Hong Kong and Alibaba's Qwen team, tackles exactly this long-standing pain point: making VLA models not just complete tasks, but complete them the way a human specifies.
The best mixed policy reaches 86.8%/82.5% in RoboTwin simulation (+15.0/+11.1 over baseline) and 62.7/100 on a real dual-arm robot (Raw-only: 49.9). Code, models, and evaluation benchmark are all open-sourced.

Contact region, target object, executing arm, trajectory direction, failure recovery — all of these execution-sensitive factors are controllable through language.
Background: Why VLA Models Still Aren't "Obedient" Enough
VLA (Vision-Language-Action) models can already grasp and place objects from natural language, but language supervision is too coarse-grained. For the same "pick up the spoon," different trajectories might use the left arm or the right arm, detour around obstacles or go straight, contact the spoon's handle or its head — yet the dataset typically shares a single goal-level instruction for all of them.
This creates supervision ambiguity: the model learns "success," but not the execution constraints of which hand to use or which direction to approach from. Building controllable VLA systems faces three core challenges:
- No infrastructure for going from heterogeneous data to fine-grained annotation
- No benchmark for evaluating robots' fine-grained understanding, and no scalable low-cost annotator
- No systematic evidence that fine-grained language actually improves policy learning
FineVLA takes on all three problems at once, forming a complete "data—model—evaluation—policy" loop.

Left: FineVLA-Tool data construction; right: FineVLA-Policy policy training.
Four Core Components
FineVLA-Tool: 970K Trajectories → 47159 Representative Samples
Converting heterogeneous robot data into high-quality fine-grained supervision, in four stages:
- Format unification: aggregate 972247 trajectories from 10 open-source datasets including Bridge V2, BC-Z, RT-1, and RoboMIND, unified into the LeRobot2.1 format
- Action normalization: unify disparate time references and kinematic representations into absolute coordinates + normalized quaternion rotations, and remove corrupted trajectories
- DTW clustering and deduplication: compute action similarity via dynamic time warping and hierarchical clustering, filtering the 970K trajectories down to 47159 representative samples
- Ten-dimensional fine-grained annotation: action sequences, executor (left/right arm), target object, contact and approach style, trajectory direction, failure recovery, and more across 10 dimensions — average word count rises from 9.3 to 96.8 (10.4x)

970K trajectories aggregated from 10 open-source datasets, then format unification + clustering and dedup + ten-dimensional fine-grained annotation.
RoboFine-VLM: Teaching a VLM to Describe How a Robot "Moves"
General VLMs often miss execution details like object disambiguation, contact regions, and motion paths. The team applied full-parameter supervised fine-tuning to Qwen3.5-VL-397B-A17B to produce RoboFine-VLM — a model that outputs step-level action descriptions covering all 10 control dimensions, and can serve as a scalable annotator for future data expansion.
RoboFine-Bench: Evaluating Fine-Grained Action Understanding
It contains 500 videos, 32 robot embodiments, and 11631 atomic facts, strictly disjoint from the training set, with two tracks:
- VQA track: 1030 questions distributed along ten dimensions, aggregated into three evaluation axes — Grounding / Action / State
- Caption track: the model must generate action-aligned step-level descriptions; an LLM judges the alignment between output and atomic facts, producing consistency, coverage, and anti-hallucination metrics
FineVLA-Policy: Verifying the Policy Gains of Fine-Grained Language
Visual observations and action labels stay unchanged; only the paired language changes (Raw-only / FG-only / Mixed), strictly isolating the effect of language supervision.
Experimental Results
Model Understanding
RoboFine-VLM scored 68.2% on the VQA track, beating the strongest general baseline GPT-5.4 (60.2%) by +8.0 points; on the Caption hard setting it reached 82.2% vs. GPT-5.4's 78.0%. Automatic scoring aligned closely with human rankings (Spearman 0.943).
RoboTwin Simulation: An Inverted-U Curve
Two key findings:
- FG-only outperforms Raw-only in every setting (gains of +1.4 to +8.1) — fine-grained supervision does not hurt task success rates
- Success rates follow an inverted-U trend, peaking between FG:Raw = 1:2 and 1:1
The best setting reaches 86.8%/82.5%, +15.0/+11.1 over baseline. Raw tells the model "what to do," FG tells it "how to do it" — the two are complementary.

Real Dual-Arm: Across-the-Board Gains in Controllable Factors
On the CobotMagic dual-arm platform, using "paired evaluation" — changing exactly one language control factor per visual scene. FG:Raw=1:1 reached 62.7/100 on Avg(ID) (Raw-only: 49.9; FG-only: 54.4).
Gains on specific control factors: pose +23, color +18, approach direction +18, rotation direction +10, executing arm +4. The larger gains concentrate precisely on the factors that goal-level instructions leave unspecified — which is exactly where FineVLA is most valuable.

Pose rises from 24 to 47, color from 22 to 40, approach direction from 60 to 78.
How to Use It
The open-source resources are complete:
- Paper: arxiv.org/abs/2605.27284
- Project page: finevla.xlang.ai
- GitHub: github.com/xlang-ai/FineVLA
- Benchmark:
RoboFine-benchon HuggingFace - Annotator model:
RoboFine-VLM-397B-A17Bon HuggingFace
Use Cases
- Robot policy research teams: reuse FineVLA-Data's 47159 fine-grained annotations directly, or use RoboFine-VLM to scale annotation of your own data
- VLA model evaluation: RoboFine-Bench provides systematic Grounding / Action / State evaluation, filling the long-standing "fine-grained understanding" gap
- Dual-arm / humanoid robot teams: the mixed-training recipe (FG:Raw=1:1) is validated to improve key control factors like pose and approach direction
- VLM for Robotics: RoboFine-VLM works as a scalable annotator, turning more unlabeled robot video into fine-grained supervision
💡 Tip: The core conclusion is not "add longer descriptions to your data" — it's that fine-grained language should augment, not replace, goal-level instructions. The best ratio is FG:Raw ≈ 1:2 to 1:1; pure fine-grained actually underperforms mixed training. Raw handles "what to do" and FG handles "how to do it" — neither works without the other.
Project page: github.com/xlang-ai/FineVLA
Toolin Editorial Team
Categories
Related articles

PerfEvolve: Teaching Agents to Tune Databases Like a Senior DBA
ISCAS open-sources the PerfEvolve framework, converting static tuning docs into executable procedural skills for Agents and delivering up to 58.9% performance gains on PostgreSQL v16 — a direct fix for LLM tuning failures.

Spatial-TTT: An Open-Source Spatial Intelligence Model at 2B Parameters
Tsinghua's open-source Spatial-TTT makes ECCV 2026: at just 2B parameters it beats GPT-5 and Gemini-3-pro on multiple spatial intelligence benchmarks, handling 120-minute streaming video while updating its spatial memory as it watches.

TerminalWorld: The First Agent Benchmark Built on Real CLI Workflows
A 1,530-task benchmark distilled from 80,000 human terminal recordings, spanning 18 workflow categories and 1,280 command tools — a cure for Agents that rack up leaderboard points yet fall apart in real terminal scenarios.

Ruoyu Lanyue 01: The World's First AI Explosion-Proof Robot Fuels Real Cars by Itself
Driven by the Ruoyu Jiutian robot brain, the explosion-proof Lanyue 01 robot autonomously runs the full workflow at gas stations 24/7 — opening the fuel door, grabbing the nozzle, fueling, returning the nozzle — bringing embodied intelligence into high-risk environments.

TRIAD: Teaching AI Agents Not Just to Refuse, but to Repair Dangerous Plans
The open-source Agent safety framework TRIAD replaces binary guardrails with three-way decisions (proceed/update/refuse), preserving the user's original task even under prompt injection attacks.

Volcano Engine's Full Agent Infra Upgrade: The 1+N+X System, AgentKit, and ArkClaw Enterprise
From the Seedance 2.0 moment to enterprise-grade Agent infrastructure, Volcano Engine uses the 1+N+X system, AgentKit, and ArkClaw Enterprise to move Agents from personal tools into organizational workflows for real.