UniRL: Tencent Hunyuan Open-Sources Its Multimodal RL Training Framework

·Toolin Editorial Team

Tencent Hunyuan's Pang Tianyu team has open-sourced UniRL, a single framework covering reinforcement learning training for image, video, and LLM multimodal generative models, supporting mainstream models like SD3, HunyuanVideo, and Qwen.

UniRL: Tencent Hunyuan Open-Sources Its Multimodal RL Training Framework

UniRL is a reinforcement learning post-training framework for multimodal generative models, open-sourced by Tencent Hunyuan's Pang Tianyu team. It solves a long-standing pain point: image diffusion models have one training pipeline, video generation has another standard, and VLMs and LLMs run on yet another tech stack. UniRL unifies all of them in a single framework.

What Problem It Solves

RL training for multimodal generative models faces four major challenges:

  1. Different generation processes: LLMs process discrete token sequences, while image/video generation corresponds to denoising trajectories in a continuous latent space. A single rollout of a unified model also mixes token generation with latent denoising.

  2. Hard-to-stabilize system loop: rollout, log-prob replay, and policy updates span multiple models and backends; the training side must strictly reproduce the sampling side's conditions, noise, and timesteps, or a training-inference mismatch appears.

  3. Heavier reward systems: multimodal rewards often depend on VLMs, OCR, aesthetic models, and video understanding models — not simple text rules.

  4. Heavy VRAM pressure: intermediate artifacts are high-dimensional latents, noise, timesteps, and conditioning states, which scale up rapidly with resolution and frame count in video generation.

The result is an industry status quo of "one model, one set of training code", with developers wasting huge amounts of time on repeated engineering work.

UniRL framework architecture

UniRL's unified multimodal RL loop: rollout -> reward -> advantage -> train -> weight-sync.

Supported Models

UniRL offers the industry's broadest support for multimodal generative models:

DomainSupported models
Image generationSD3/3.5, Qwen-Image, Z-Image, FLUX.2-Klein
Video generationHunyuanVideo 1.0&1.5, WAN series
Large language modelsQwen3 series
Multimodal understandingQwen-VL series
Natively unified modelsHunyuanImage 3.0, Bagel
Composable modelsLLM/VLM + Diffusion Prompt-Enhancer

Built-in Algorithms and Reward Systems

Algorithm Side

  • Policy-gradient family: FlowGRPO, DanceGRPO, MixGRPO, LLM/VLM GRPO
  • Forward-process family: DiffusionNFT (efficient training without a full SDE rollout)
  • Tencent Hunyuan in-house:
    • Flow-DPPO: replaces PPO ratio clipping with a per-step KL-divergence proximal constraint, achieving more stable RL training of flow/diffusion models
    • DRPO: replaces hard clipping/masking with an advantage-weighted smooth policy-shift regularizer, providing continuous gradient correction even past the trust-region boundary

Reward Side

UniRL integrates many commonly used reward components:

  • Rule/similarity: CLIPScore, GOT-OCR-2.0
  • Preference/aesthetics: PickScore, HPSv2/HPSv3, ImageReward
  • VLM-as-judge: UnifiedReward, GenEval2, WISE
  • Video evaluation: VideoPickScore, VideoAlign

Core Design

UniRL is built around Ray worker groups, a Hydra flat recipe, composable training backends, and pluggable rollout engines as its core skeleton. It uses tracks to represent the generation trajectories carried through different stages: the AR stage is a TextSegment, the image generation stage is a LatentSegment, and different tracks connect through parent-child relationships.

This lets the chained flow of unified multimodal models like Bagel and HunyuanImage 3.0 (AR text thinking first, then DiT image generation) be represented naturally.

How to Use It

The framework ships with comprehensive examples, making it easy to quickly launch experiments and reproduce algorithms. It is still under active iteration; rollout engine support will continue to expand and large-scale training performance will keep improving.

Related articles

TerminalWorld: The First Agent Benchmark Built on Real CLI Workflows
AI Tutorials

TerminalWorld: The First Agent Benchmark Built on Real CLI Workflows

A 1,530-task benchmark distilled from 80,000 human terminal recordings, spanning 18 workflow categories and 1,280 command tools — a cure for Agents that rack up leaderboard points yet fall apart in real terminal scenarios.

Toolin Editorial Team
Ruoyu Lanyue 01: The World's First AI Explosion-Proof Robot Fuels Real Cars by Itself
AI Products

Ruoyu Lanyue 01: The World's First AI Explosion-Proof Robot Fuels Real Cars by Itself

Driven by the Ruoyu Jiutian robot brain, the explosion-proof Lanyue 01 robot autonomously runs the full workflow at gas stations 24/7 — opening the fuel door, grabbing the nozzle, fueling, returning the nozzle — bringing embodied intelligence into high-risk environments.

Toolin Editorial Team
TRIAD: Teaching AI Agents Not Just to Refuse, but to Repair Dangerous Plans
AI Products

TRIAD: Teaching AI Agents Not Just to Refuse, but to Repair Dangerous Plans

The open-source Agent safety framework TRIAD replaces binary guardrails with three-way decisions (proceed/update/refuse), preserving the user's original task even under prompt injection attacks.

Toolin Editorial Team
Volcano Engine's Full Agent Infra Upgrade: The 1+N+X System, AgentKit, and ArkClaw Enterprise
AI Products

Volcano Engine's Full Agent Infra Upgrade: The 1+N+X System, AgentKit, and ArkClaw Enterprise

From the Seedance 2.0 moment to enterprise-grade Agent infrastructure, Volcano Engine uses the 1+N+X system, AgentKit, and ArkClaw Enterprise to move Agents from personal tools into organizational workflows for real.

Toolin Editorial Team
Baidu DuMate Hands-On Guide: A Homegrown Codex That Lets You Run Office Work by Voice
AI Tutorials

Baidu DuMate Hands-On Guide: A Homegrown Codex That Lets You Run Office Work by Voice

From installation to automation, a complete breakdown of Baidu DuMate's request -> authorize -> execute -> deliver pipeline: cross-app work, scheduled tasks, and the credits bill, all explained in one article.

Toolin Editorial Team
DeNovoSWE: The First Long-Horizon Doc2Repo Training Set That Teaches Code Agents to Build Repositories
AI Products

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set That Teaches Code Agents to Build Repositories

Renmin University of China releases an open training set of 4818 real task instances targeting repository-level code generation, lifting Qwen3-30B's pass rate from 5.8% to 47.2%.

Toolin Editorial Team