FineVLA Goes Open Source: One Sentence Tells the Robot Which Hand to Use and Where to Grab

·Toolin Editorial Team

HKU and Alibaba jointly open-source FineVLA, a controllable VLA framework that lets language specify the executing arm, contact region, and other details, with an 86.8% success rate in RoboTwin simulation.

FineVLA Goes Open Source: One Sentence Tells the Robot Which Hand to Use and Where to Grab

Today's robot models can understand "put the cup in the basket" — but which hand? From which direction? Grab the body or the handle? These key details determine execution quality, yet existing datasets rarely label them. The XLANG Lab at the University of Hong Kong and Alibaba's Qwen team have jointly open-sourced FineVLA — an open-source framework for controllable VLA (Vision-Language-Action) policies that lets robots not only complete a task but complete it the way a human specifies. Code, models, and the evaluation benchmark are all open-sourced.

What FineVLA Is

Think of it as "the language layer that lets robots understand execution details." In image generation, the details of a text description directly affect how controllable the result is; robot policy learning works the same way — language needs to constrain the actual action process. Two trajectories for picking up the same spoon might use the left arm or the right arm, skirt an obstacle or move in a straight line, yet in a dataset they often share the same goal-level instruction.

This creates supervision ambiguity: the model can learn "the task must eventually succeed," but it struggles to learn from language the execution constraints — which hand to use, which direction to approach from, which part of the object to touch. FineVLA was born to fill in exactly this layer.

FineVLA turns "put the cup in the basket" into an executable, concrete instruction like "use the left hand, approach from the right, grab the handle."

Core Components: Four Modules Forming a Closed Loop

FineVLA builds a complete "data — model — evaluation — policy" loop.

On the left is the data construction pipeline (10 datasets → 970,000 trajectories → 47,000 representative samples → ten-dimension annotation); on the right are policy learning and evaluation.

Component One: FineVLA-Tool — From 970,000 Trajectories to Fine-Grained Data

Through four stages, heterogeneous robot data is converted into high-quality fine-grained supervision:

  • Stage 1, format unification: aggregate 972,247 trajectories from 10 open-source datasets including Bridge V2, BC-Z, RT-1, and RoboMIND, converting them uniformly into the LeRobot2.1 format
  • Stage 2, action canonicalization: unify temporal reference and kinematic representation into absolute coordinates plus normalized quaternion rotations, and remove corrupted trajectories
  • Stage 3, DTW clustering and deduplication: compute action-trajectory similarity via dynamic time warping and hierarchical clustering, filtering 970,000 trajectories down to 47,159 representative samples
  • Stage 4, ten-dimension fine-grained annotation: annotate along 10 dimensions such as action sequence, executor (left/right arm), target object, contact and approach style, trajectory direction, and failure recovery. Average word count rises from 9.3 to 96.8 after annotation (a 10.4x increase)

Component Two: RoboFine-VLM — Teaching a VLM to Describe How a Robot Moves

General-purpose VLMs often miss execution details like disambiguating similar objects, contact regions, and motion paths. The team performed full-parameter supervised fine-tuning on Qwen3.5-VL-397B-A17B to produce RoboFine-VLM, which outputs step-level action descriptions covering all 10 control dimensions and serves as a scalable annotator for future data expansion.

Component Three: RoboFine-Bench — Evaluating Fine-Grained Action Understanding

The benchmark strictly does not overlap with the training set; it has VQA and Caption tracks covering three evaluation axes: localization, action understanding, and state reasoning.

It contains 500 videos, 32 robot embodiments, and 11,631 atomic facts, with strict non-overlap against the training set. There are two tracks:

  • VQA track: 1030 questions distributed across the ten fine-grained dimensions
  • Caption track: the model must generate action-aligned step-level descriptions, judged by an LLM on consistency, coverage, and anti-hallucination

Component Four: FineVLA-Policy — Verifying the Policy Gains from Fine-Grained Language

Three configurations are designed to strictly isolate the effects of "architecture" versus "data scale," and each configuration is evaluated under seven FG:Raw instruction ratios.

Hands-On: The Results Data

Simulation and Real-Robot Results

Scores for the best mixed-policy setting:

Test environmentMetricFineVLABaselineImprovement
RoboTwin simulationSuccess rate86.8% / 82.5%Baseline+15.0 / +11.1
Real dual-arm robotScore62.7 / 100Raw-only 49.9+12.8

Broken down by controllable factor, dimensions such as pose (+23), color (+18), and approach direction (+18) all improve significantly.

Strengths

  • Genuinely controllable: no longer "can complete the task" but "completes it the way you said"
  • Complete data loop: from heterogeneous data to fine-grained annotation to evaluation to policy, every link is open-sourced
  • Publicly available benchmark: RoboFine-Bench fills a gap in fine-grained robot understanding evaluation

Limitations

  • Validated mainly in dual-arm settings: transfer to single arms, mobile manipulation, and other embodiments needs your own testing
  • Depends on annotation quality: fine-grained annotations are currently generated by Qwen3.5-Plus plus human review; scaling up depends on the stability of RoboFine-VLM

Use Cases

  • Embodied AI research: open code + models + benchmark make it a directly usable research baseline for controllable VLA
  • Industrial dual-arm robots: scenarios that need precise control of execution details (which hand, where to grab)
  • Robot data annotation: as a scalable annotator, RoboFine-VLM can accelerate fine-grained annotation of new datasets
  • VLA model evaluation: RoboFine-Bench is a public yardstick for how "obedient" a model is

Code, models, and the evaluation benchmark are all open-sourced on GitHub (search for "FineVLA").