Jetson-PI: Peking University Open-Sources Real-Time On-Device VLA Control, 8.66× Higher Control Frequency on Jetson Orin
Peking University, AIRS, and PrimeBot open-source Jetson-PI (Apache-2.0): FAAC asynchronous inference + confidence scheduling + a llama.cpp engine lift π0.5 on Jetson Orin from 0.70Hz to 6.06Hz, with no loss in accuracy.


Jetson-PI: Peking University Open-Sources Real-Time On-Device VLA Control, 8.66× Higher Control Frequency on Jetson Orin
Peking University, AIRS, and PrimeBot open-source Jetson-PI (Apache-2.0): FAAC asynchronous inference + confidence scheduling + a llama.cpp engine lift π0.5 on Jetson Orin from 0.70Hz to 6.06Hz, with no loss in accuracy.
Getting a VLA (Vision-Language-Action) model off an RTX 4090 workstation and onto the robot's own body has long been the hard nut of embodied AI deployment. π0.5 takes roughly 1.4 seconds per inference on an NVIDIA Jetson Orin, which works out to a control frequency of just 0.7 Hz — sluggish robot motion, with visible stalls between consecutive action chunks. Jetson-PI (Apache-2.0), open-sourced jointly by Peking University, AIRS, and PrimeBot Research Institute, delivers a complete answer: no quantization, no pruning — instead it attacks the problem at three layers at once (asynchronous inference algorithm, model scheduling, and the underlying inference engine), lifting π0.5's control frequency on Jetson Orin to 6.06 Hz (8.66×) with no loss in accuracy. It's a key step from "lab demo" toward "industrial-grade, runs for the long haul."
Paper: https://arxiv.org/abs/2607.12659 Asynchronous inference code: https://github.com/PKU-SEC-Lab/Jetson-PI On-device inference engine: https://github.com/PKU-SEC-Lab/Jetson-PI-Edge
The Problem: On-Device VLA Isn't Just "Slow Inference"
Under traditional synchronous inference, the robot must wait for the model to finish generating a complete action chunk before it can start executing; on-device compute is limited, so inference latency shows up directly as long stalls. The natural fix is asynchronous inference: while the current action chunk executes, the model predicts the next chunk in parallel. But async inference introduces two new problems:
- Perception–execution misalignment: the model predicts the next action chunk using the image captured when inference started, but by the time the action actually executes, the robot and the environment have already changed — the longer the delay, the bigger the error.
- Reaction time is still too long: even with parallelism, the robot's response to environmental changes (an object slipping, a position shifting) remains bounded by the full VLA inference latency, and a single on-device inference often exceeds 1 second.
Past asynchronous methods tried blending old and new trajectories along the time axis to improve smoothness, but action prediction itself still suffered from the two problems above. Jetson-PI takes a different road: don't blend trajectories — directly predict the future environment representation, and start predicting actions from that future moment.
Core Algorithm: FAAC (Foresight-Aligned Asynchronous Correction)
Foresight-Aligned Asynchronous Correction (FAAC) is the heart of Jetson-PI's method. The team trained a lightweight Future Correction Module:
- Input: the environment representation produced by the VLM at the current moment + the action sequence already committed to the robot and about to execute
- Output: the VLM environment representation of the environment at a future moment, after those actions finish executing
- The action expert no longer starts from the stale current observation; it generates action sequences beginning at the projected future point in time
Restraint in the key design:
- It does not generate full future images (far too expensive computationally)
- It does not correct the VLM's entire KV Cache layer by layer (which would introduce a large amount of extra inference)
- It only compresses and predicts the VLM's final-layer environment representation, feeding the correction into the action expert
The whole future correction module is about 40M parameters, only around 1% of the full VLA model's parameter count, so it puts little burden on on-device hardware. During training, the future time horizon is randomly sampled, letting the same correction module adapt to the different latencies of a Jetson Orin, a Jetson Thor, or a more powerful GPU.
Confidence Scheduling: Teaching the Robot "When to Take Another Look"
With future correction in place, in theory one VLM observation should suffice — keep calling the future correction module and the action expert, and actions keep flowing. But future correction carries error, and relying on it alone accumulates mistakes over time.
Jetson-PI's insight is to rethink the division of labor between the VLM and the action expert:
- The VLM's main job is to re-observe the environment, correcting the understanding of the external world
- The action expert handles high-frequency action generation, with faster inference
Hence the confidence-based scheduling mechanism: the future correction module outputs a confidence score alongside its predicted future representation.
- Confidence above threshold → skip the VLM and call the action expert directly on the corrected representation (the robot generates new actions more frequently)
- Confidence below threshold → re-invoke the VLM to refresh the environment representation, the KV Cache, and the state caches later predictions depend on
In practice, confidence is usually high during smooth-motion phases where environmental change is easy to predict; during critical steps like grasping, contact, and placement it drops noticeably, triggering a fresh VLM observation. It amounts to teaching the robot to judge "when I can keep going on what I already know, and when I must look again."
Jetson-PI-Edge: Rebuilding the On-Device Inference Engine on llama.cpp
The algorithm answers "when to call the model," but for π0 / π0.5 to truly hit real-time on-device control, the underlying inference system needs rebuilding too. The team built Jetson-PI-Edge on llama.cpp, with three core optimizations tailored to how VLA inference behaves:
1. Graph Reuse
Traditional language models produce output whose length changes dynamically as decoding proceeds, so the computation graph's shape keeps changing. VLA inference is different: with camera count, image resolution, and action dimensions fixed, most input tensor sizes are fixed; language instructions are usually short and can be padded to a fixed length. So the CUDA Graph is built on the first inference and reused directly afterward, avoiding graph construction repeated on every control round.
2. GPU-Resident Intermediate Buffers
ViT / LLM / action expert pass intermediate results — visual embeddings, KV Cache, and so on — between one another. Generic inference frameworks write them back to CPU memory, and the next module copies them onto the GPU again; on bandwidth-limited on-device hardware this H2D / D2H round-trip overhead is significant. Jetson-PI-Edge pre-allocates GPU buffers for fixed-size intermediate results, letting ViT, LLM, and the action expert reuse the data directly on the GPU.
3. Flow Unrolling
The π-family action experts run multiple rounds of flow-matching denoising. Invoking the graph separately at every step means repeated scheduling and kernel-launch overhead. Jetson-PI-Edge unrolls the denoising rounds into one unified computation graph — graph construction costs a bit more the first time, but stays continuously reusable afterward, which suits long-running robot control loops especially well.
Benchmark: 8.66× Speedup, Top Average SR on LIBERO
On-Device Latency (NVIDIA Jetson Orin, MAXN, ms)
| Stage | PI0.5 total latency | PI0 total latency |
|---|---|---|
| Naive | 1420.8 (ViT 152.3 / LLM 631.0 / AE 536.8) | 1250.9 |
| +Schedule optimization | 1420.8 | 1251.5 |
| +Graph reuse | 476.1 | 444.4 |
| +Buffer + Unroll | 412.9 (ViT 79.5 / LLM 210.3 / AE 123.1) | 394.5 (ViT 75.4 / LLM 200.3 / AE 118.8) |
With confidence scheduling layered on, system reaction time drops further to 165.1 ms and control frequency reaches 6.06 Hz, an 8.66× improvement over native PyTorch (0.70 Hz). Note that this speedup does not depend on quantization or pruning, and in theory it can be combined with quantization and pruning to push latency even lower.
LIBERO (π0.5) Average Success Rate Across Four Subsets
Jetson-PI (Ours+Sched) ranks first on all four subsets:
| Subset | Jetson-PI SR |
|---|---|
| SPATIAL | 97.4 |
| OBJECT | 98.6 |
| GOAL | 96.8 |
| LIBERO-10 | 92.5 |
As asynchronous latency grows, VLASH, which only predicts the robot's future state, degrades sharply. At Δ=9, Jetson-PI beats VLASH by 45.6 percentage points and RTC by 7.0 percentage points on the four-subset average; overall it averages 14.8 pp above VLASH and 3.9 pp above RTC. These experiments make the point: on-device VLA isn't just "making inference fast" — the model also has to understand how the actions the robot executes during the latency window will change the future environment.
Real-Robot Experiments
On the PrimeBot X2-W robot (1 head + 2 wrist cameras, 224×224, 15 Hz action execution), tasks are split into three subtasks: garment pickup / unfold-and-fold / tidying away. Jetson-PI's trajectories are more continuous, garment manipulation is smoother, and critical phases no longer fail from perception–execution misalignment.
Getting Started on the Training Side: Jetson-PI (JAX/Python 3.11)
Environment Requirements
- OS: Ubuntu 22.04
- GPU: NVIDIA GPU with ≥ 48 GB VRAM (for batch-16 three-stage full training)
- CUDA: 12.x (installed as a project dependency; no system CUDA needed)
- Python: 3.11 (uv / JAX)
Installation
git clone --recurse-submodules https://github.com/PKU-SEC-Lab/Jetson-PI
cd Jetson-PI
git submodule update --init --recursive
export PYTHONNOUSERSITE=1
GIT_LFS_SKIP_SMUDGE=1 uv sync
GIT_LFS_SKIP_SMUDGE=1 uv pip install -e .
# Apply the π0.5 PyTorch/JAX compatibility patch
cp -r ./src/openpi/models_pytorch/transformers_replace/* \
.venv/lib/python3.11/site-packages/transformers/Downloading Weights (ModelScope)
The weights repo zebinyang/Jetson-PI-pi05 contains two directories, pi05_libero/ and future_correction_module/; do not merge the params trees.
pip install modelscope
python -c "from modelscope import snapshot_download; snapshot_download('zebinyang/Jetson-PI-pi05', local_dir='./checkpoints/jetson-pi-pi05')"
export PI0_CHECKPOINT=./checkpoints/jetson-pi-pi05/pi05_libero
export WM=./checkpoints/jetson-pi-pi05/future_correction_moduleDefault Training Recipe (π0.5-LIBERO, Three Stages)
| Stage | Steps | What is trained |
|---|---|---|
| Stage 1 | 30000 | Action Expert + token reducer (L_act) |
| Stage 2 | 15000 | Future correction module (L_cond, no logvar head) |
| Stage 3 | 55000 | L_cond (no reducer) + L_act on Pi0 AE + full LLM (μ detached) |
Handover settings are fixed at H=10, max_delta_t=10, action_encoder=transformer_block.
bash scripts/train_wm_libero_spatial_four_stage.shYou can override STAGE1_STEPS / STAGE2_STEPS / STAGE3_STEPS / BATCH_SIZE / NUM_WORKERS / EXP_NAME; logs land in logs/<EXP_NAME>.log.
Single Evaluation Run
export PI0_CHECKPOINT=PATH/TO/CHECKPOINT/pi05_libero
export PY_SERVER=PATH/TO/PYTHON # must be a venv with JAX
export WM=PATH/TO/future-correction-module
export CUDA_VISIBLE_DEVICES=0
export PORT=8000
bash scripts/eval_wm_libero_spatial.shDefaults to libero_spatial, 50 trials per task, H=10, K=9, overlap=1.
Confidence scheduling (adaptive multi-rollout):
export LIBERO_WM_EVAL_ADAPTIVE_KAPPA=1
export LIBERO_WM_EVAL_KAPPA_DELTA=0.4
bash scripts/eval_wm_libero_spatial.shTo switch to another task suite: export LIBERO_WM_EVAL_TASK_SUITE=libero_object|libero_goal|libero_10.
💡 Pitfalls:
ModuleNotFoundError: jax→ pointPY_SERVERat a venv that has JAX- OOM → lower
BATCH_SIZE/ setXLA_PYTHON_CLIENT_MEM_FRACTION=0.85/ setNUM_WORKERS=0- Missing
norm_stats→ point--pi0-norm-checkpoint-dirat the tree containingassets/physical-intelligence/libero/norm_stats.json- EGL/display issues → install
xvfbor useMUJOCO_GL=egl
Getting Started on the Edge Side: Jetson-PI-Edge (llama.cpp)
Build
CPU-only:
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-server -jJetson / CUDA:
cmake -S . -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-server -jRunning the Foreground Server
PI_MODEL supports auto | pi0 | pi05 (the default auto is detected from GGUF metadata and tensor names):
PI_MODEL=auto ./build/bin/llama-server \
-m /path/to/pi_llm.gguf --mmproj /path/to/mmproj.gguf \
-ngl 99 --host 0.0.0.0 --port 8080Single Inference over HTTP
Four steps, in order:
POST /foreground/reset— initialize the sessionPOST /foreground/image(twice, for the two camera views)PUT /foreground/state(32-dim robot state as CSV)POST /foreground/inferwith body{"text":"pick up the object and place it into the tray"}
The response includes action_final plus encode_ms / decode_ms / total_ms / timing_breakdown_ms, making it easy to localize latency module by module.
Python Foreground Client (Reusing a Session Across Control Steps)
from jetson_pi_foreground import ManagedForegroundSession
session = ManagedForegroundSession(
server_path=...,
model_path=...,
mmproj_path=...,
gpu=0,
port=8080,
timeout=300,
)
action, metadata = session.predict(
image_paths=[IMG, IMG],
prompt='/do something',
state=state_np_float32,
reset=True,
)
session.close()Reusing the same session avoids rebuilding the CUDA context on every step — a critical performance saving inside the control loop.
FlashRT Integration (Optional)
Through a C API provider, the same GGUF runtime can be exposed to FlashRT's Python interface without running the foreground HTTP server. Configure cmake with -DFLASHRT_CPP_WITH_JETSON_PI=ON -DJETSON_PI_ROOT=/path/to/Jetson-PI-Edge -DGGML_CUDA=ON -DGGML_CUDA_FA=ON and build libflashrt_cpp_llama_cpp_provider_c.so.
Roadmap Models
The Jetson-PI framework isn't locked to π0 / π0.5; upcoming support includes:
- NVIDIA Isaac GR00T N1.7
- LingBot-VLA 2.0
- Qwen-RobotManip
- DreamZero
- FastWAM
Use Cases and Limits
A good fit for:
- Mobile robotics projects deploying VLA on on-device hardware like the Jetson Orin / Thor, where power draw and battery life matter
- Embodied-AI researchers who want to run the full π0.5 training + evaluation pipeline end to end
- Engineering teams that need to plug VLA inference into a real-time control loop (15Hz+)
Current limitations:
- The training-side hardware bar is high (≥48 GB VRAM)
- The on-device optimizations are tightly bound to the llama.cpp architecture; deep customization requires C++ experience
- Real-robot experiment data is still concentrated on garment-folding tasks; more task types await community validation
Final Thoughts
What Jetson-PI cares about is not pushing VLA parameters ever larger, but solving a problem closer to real robot deployment: when a model has to run on an onboard device constrained in power, bandwidth, and thermals, how do you keep its reaction speed fast enough. The answer has three parts: use future environment representations to mitigate perception–execution misalignment in asynchronous inference; use confidence scheduling to cut unnecessary VLM calls; use VLA-specific system optimizations to compress on-device execution overhead. Code, engine, weights, and training scripts are all open-sourced — for developers working on embodied AI, this is a 2026 project worth forking outright as a starting point.
Paper: https://arxiv.org/abs/2607.12659 Asynchronous inference code: https://github.com/PKU-SEC-Lab/Jetson-PI On-device inference engine: https://github.com/PKU-SEC-Lab/Jetson-PI-Edge