DyRef Hands-On: Adding Multi-Reference Consistency to Qwen Image Editing Models—Stable Even at 7 Reference Images
DyRef from HIT Shenzhen (ECCV 2026 Oral) is a multi-reference consistency fine-tuning recipe + RL dynamic-reward training framework layered on Qwen-Image-Edit-2511, keeping subject, pose, and style consistent even with up to 7 reference images. This guide runs local inference, plus optional SFT/RL training.

DyRef is an image-editing training framework from Zhu Taotao's research group at Harbin Institute of Technology, Shenzhen (first author Wenwang Huang et al.). The paper, "Scaling Multi-Reference Image Generation with Dynamic Reward Optimization" (arXiv 2606.26947), has been accepted as an ECCV 2026 Oral (announced 2026-07-24).
Its angle of attack is very specific: existing open-source image editing models (such as Qwen-Image-Edit-2511 and FLUX.2-klein-base-9B) do reasonably well when you give them a single reference image, but the moment you feed in multiple references playing different roles (subject / background / pose / lighting / style), you get "the person changed, the pose doesn't line up, the style drifted." DyRef exists to patch exactly this—it can handle roughly 7 reference images in different roles at the same time.
Be clear about what it is and isn't: on 2026-07-21 we published qwen-image-3 (Alibaba Qwen's third-generation image generation base model). DyRef is not a new base model. It is a "multi-reference consistency fine-tuning recipe + RL dynamic-reward training framework" layered on top of Qwen-Image-Edit-2511—think of it as a "multi-reference constraint plugin" installed on an open-source image editing model. That's the reason it merits its own article.
The Problem It Solves
Feed multiple reference images to an open-source image editing model and the usual failure modes look like this:
- Subject drift: you specify person A's face plus B's background, and the generated face comes out different
- Pose misalignment: you supply a pose reference, and the final pose is off by 30 degrees
- Style bleed: you specify cyberpunk, and the output leans photorealistic
DyRef uses two-stage training (SFT + RL) to keep the model consistent under multi-reference constraints, and on the authors' OmniRef-Bench (395 expert-annotated samples covering 5 reference types × 10 combinations) it matches closed-source Nano Banana Pro and surpasses Seedream 4.5.
Open-Source Artifacts (all Apache-2.0, individually verified)
| Artifact | URL | Notes |
|---|---|---|
| Code repo | github.com/Weistrass/DyRef | Contains sft/, rl/, model_inference/, benchmark/; 318 stars |
| Model weights | huggingface.co/Weistrass/Qwen-Image-Edit-2511-DyRef | LoRA weights |
| Training data | huggingface.co/datasets/Eason0438/OmniRef-training | ~14K samples suffice for training |
| Benchmark | huggingface.co/datasets/Eason0438/OmniRef-Bench | 395 multi-reference samples |
| Paper | arxiv.org/abs/2606.26947 | ECCV 2026 Oral |
💡 Tip: data efficiency is high—the full SFT+RL pipeline runs on only about 14K samples, friendly to memory and time budgets.
Before You Start
- Hardware: an NVIDIA GPU (PyTorch + CUDA environment); for training, 24GB+ of VRAM recommended
- Accounts: GitHub (pull code) + Hugging Face (pull weights and datasets)
- Base model: Qwen-Image-Edit-2511 (FLUX.2-klein-base-9B also supported)
- Expected time: inference only, about 10 minutes; full SFT+RL training depends on data size
Step 1: Create the conda environments
The repo ships separate environment configs for SFT and RL (based on Python 3.11); choose according to your goal.
Just run inference with the released weights (recommended for beginners)—only the SFT environment is needed:
conda create -n dyref_sft python=3.11 -y
conda activate dyref_sft
cd DyRef/sft
pip install -e .To run the full SFT + RL training—install both environments:
# SFT environment (same as above)
conda create -n dyref_sft python=3.11 -y && conda activate dyref_sft && cd DyRef/sft && pip install -e .
# RL environment (adds deepspeed)
conda create -n dyref_rl python=3.11 -y && conda activate dyref_rl && cd DyRef/rl && pip install -e .[deepspeed]💡 Tip: the RL environment installs with
pip install -e .[deepspeed]because Stage 2 training relies on DeepSpeed for distributed optimization.
Step 2: Pull the LoRA weights
Pull the officially released LoRA weights from Hugging Face into the expected local path (the README downloads from the HF Hub automatically; users in mainland China should download them manually first):
# Weights URL
# https://huggingface.co/Weistrass/Qwen-Image-Edit-2511-DyRefThis is a LoRA adapter stacked on the Qwen-Image-Edit-2511 base—far smaller than full model weights.
Step 3: Run multi-reference inference (fastest start)
If you only want to try "multi-reference image editing" and don't need to train, run the official inference script directly:
conda activate dyref_sft
python model_inference/Qwen-Image-Edit-2511.pyThis script loads base + LoRA weights, accepts multiple reference images (subject/background/pose/lighting/style roles), and outputs one composite that fuses all the constraints. Start by picking a few combinations from OmniRef-Bench's 395 samples to verify the environment works.
Step 4 (optional): Stage 1 SFT supervised fine-tuning
To fine-tune on your own multi-reference data, enter the SFT stage. The official training script is ready to run out of the box:
cd DyRef/sft
bash all_scripts/Qwen-Image-Edit-2511_lora.shThis stage does LoRA fine-tuning on multi-reference data so the model first learns to "understand" the constraints of multiple reference images. The training data can be the official OmniRef-training set (~14K samples) or your own multi-reference dataset.
Step 5 (optional): Stage 2 RL reinforcement learning (the core innovation)
This is the key that separates DyRef from ordinary LoRA fine-tuning. Stage 2 optimizes further with reinforcement learning, introducing two dynamic reward mechanisms:
- DRS (Discriminative Reward Scaling)
- DAR (Difficulty-aware Advantage Reweighting)
Both attack the same pain point: in RL training, the reward signal lacks separation between easy and hard samples—easy samples score full marks, hard samples fail across the board, and the model never learns fine control over difficult multi-reference combinations. DRS + DAR "spread out" the reward signal across the difficulty gradient, so the model learns hard combinations more solidly.
Run the official RL script:
conda activate dyref_rl
cd DyRef/rl
bash scripts/qwen2511-gdpo-rank64-add2k5-csd-siglipv2_flat-sigmoid0.65-focal_loss.shThe script name is dense with information: rank64 is the LoRA rank, siglipv2 is the vision encoder used for reward, and sigmoid0.65 and focal_loss are implementation details of DAR's difficulty awareness.
Results
After training, validate multi-reference consistency on OmniRef-Bench's test set:
# Run the evaluation inside benchmark/
cd DyRef/benchmark
# The official benchmark script runs 395 samples across 5 reference types × 10 combinationsSuccess criteria: subject consistency (no face drift), pose alignment, and style fidelity all improve markedly over the base Qwen-Image-Edit-2511, reaching the paper's reported "matches Nano Banana Pro, surpasses Seedream 4.5" level.
FAQ
- Not enough VRAM: SFT uses LoRA (already the default) and is light on memory; for RL, enable DeepSpeed ZeRO-2/3 sharding. Beginners should start with step 3 inference only.
- No multi-reference data of your own: reproduce the paper's results directly with the official OmniRef-training set (~14K samples).
- Want a different base model: the repo also supports FLUX.2-klein-base-9B; see the corresponding script under
model_inference/. - The training script name is unreadable: that parameter string is actually DyRef's core hyperparameters (rank, sigmoid, focal_loss, etc.); read the RL section of the paper before changing anything.