Toolin.ai

DyRef Hands-On: Adding Multi-Reference Consistency to Qwen Image Editing Models—Stable Even at 7 Reference Images

Published · toolin小编

DyRef from HIT Shenzhen (ECCV 2026 Oral) is a multi-reference consistency fine-tuning recipe + RL dynamic-reward training framework layered on Qwen-Image-Edit-2511, keeping subject, pose, and style consistent even with up to 7 reference images. This guide runs local inference, plus optional SFT/RL training.

DyRef Hands-On: Adding Multi-Reference Consistency to Qwen Image Editing Models—Stable Even at 7 Reference Images

DyRef is an image-editing training framework from Zhu Taotao's research group at Harbin Institute of Technology, Shenzhen (first author Wenwang Huang et al.). The paper, "Scaling Multi-Reference Image Generation with Dynamic Reward Optimization" (arXiv 2606.26947), has been accepted as an ECCV 2026 Oral (announced 2026-07-24).

Its angle of attack is very specific: existing open-source image editing models (such as Qwen-Image-Edit-2511 and FLUX.2-klein-base-9B) do reasonably well when you give them a single reference image, but the moment you feed in multiple references playing different roles (subject / background / pose / lighting / style), you get "the person changed, the pose doesn't line up, the style drifted." DyRef exists to patch exactly this—it can handle roughly 7 reference images in different roles at the same time.

Be clear about what it is and isn't: on 2026-07-21 we published qwen-image-3 (Alibaba Qwen's third-generation image generation base model). DyRef is not a new base model. It is a "multi-reference consistency fine-tuning recipe + RL dynamic-reward training framework" layered on top of Qwen-Image-Edit-2511—think of it as a "multi-reference constraint plugin" installed on an open-source image editing model. That's the reason it merits its own article.

The Problem It Solves

Feed multiple reference images to an open-source image editing model and the usual failure modes look like this:

  • Subject drift: you specify person A's face plus B's background, and the generated face comes out different
  • Pose misalignment: you supply a pose reference, and the final pose is off by 30 degrees
  • Style bleed: you specify cyberpunk, and the output leans photorealistic

DyRef uses two-stage training (SFT + RL) to keep the model consistent under multi-reference constraints, and on the authors' OmniRef-Bench (395 expert-annotated samples covering 5 reference types × 10 combinations) it matches closed-source Nano Banana Pro and surpasses Seedream 4.5.

Open-Source Artifacts (all Apache-2.0, individually verified)

ArtifactURLNotes
Code repogithub.com/Weistrass/DyRefContains sft/, rl/, model_inference/, benchmark/; 318 stars
Model weightshuggingface.co/Weistrass/Qwen-Image-Edit-2511-DyRefLoRA weights
Training datahuggingface.co/datasets/Eason0438/OmniRef-training~14K samples suffice for training
Benchmarkhuggingface.co/datasets/Eason0438/OmniRef-Bench395 multi-reference samples
Paperarxiv.org/abs/2606.26947ECCV 2026 Oral

💡 Tip: data efficiency is high—the full SFT+RL pipeline runs on only about 14K samples, friendly to memory and time budgets.

Before You Start

  • Hardware: an NVIDIA GPU (PyTorch + CUDA environment); for training, 24GB+ of VRAM recommended
  • Accounts: GitHub (pull code) + Hugging Face (pull weights and datasets)
  • Base model: Qwen-Image-Edit-2511 (FLUX.2-klein-base-9B also supported)
  • Expected time: inference only, about 10 minutes; full SFT+RL training depends on data size

Step 1: Create the conda environments

The repo ships separate environment configs for SFT and RL (based on Python 3.11); choose according to your goal.

Just run inference with the released weights (recommended for beginners)—only the SFT environment is needed:

conda create -n dyref_sft python=3.11 -y
conda activate dyref_sft
cd DyRef/sft
pip install -e .

To run the full SFT + RL training—install both environments:

# SFT environment (same as above)
conda create -n dyref_sft python=3.11 -y && conda activate dyref_sft && cd DyRef/sft && pip install -e .

# RL environment (adds deepspeed)
conda create -n dyref_rl python=3.11 -y && conda activate dyref_rl && cd DyRef/rl && pip install -e .[deepspeed]

💡 Tip: the RL environment installs with pip install -e .[deepspeed] because Stage 2 training relies on DeepSpeed for distributed optimization.

Step 2: Pull the LoRA weights

Pull the officially released LoRA weights from Hugging Face into the expected local path (the README downloads from the HF Hub automatically; users in mainland China should download them manually first):

# Weights URL
# https://huggingface.co/Weistrass/Qwen-Image-Edit-2511-DyRef

This is a LoRA adapter stacked on the Qwen-Image-Edit-2511 base—far smaller than full model weights.

Step 3: Run multi-reference inference (fastest start)

If you only want to try "multi-reference image editing" and don't need to train, run the official inference script directly:

conda activate dyref_sft
python model_inference/Qwen-Image-Edit-2511.py

This script loads base + LoRA weights, accepts multiple reference images (subject/background/pose/lighting/style roles), and outputs one composite that fuses all the constraints. Start by picking a few combinations from OmniRef-Bench's 395 samples to verify the environment works.

Step 4 (optional): Stage 1 SFT supervised fine-tuning

To fine-tune on your own multi-reference data, enter the SFT stage. The official training script is ready to run out of the box:

cd DyRef/sft
bash all_scripts/Qwen-Image-Edit-2511_lora.sh

This stage does LoRA fine-tuning on multi-reference data so the model first learns to "understand" the constraints of multiple reference images. The training data can be the official OmniRef-training set (~14K samples) or your own multi-reference dataset.

Step 5 (optional): Stage 2 RL reinforcement learning (the core innovation)

This is the key that separates DyRef from ordinary LoRA fine-tuning. Stage 2 optimizes further with reinforcement learning, introducing two dynamic reward mechanisms:

  • DRS (Discriminative Reward Scaling)
  • DAR (Difficulty-aware Advantage Reweighting)

Both attack the same pain point: in RL training, the reward signal lacks separation between easy and hard samples—easy samples score full marks, hard samples fail across the board, and the model never learns fine control over difficult multi-reference combinations. DRS + DAR "spread out" the reward signal across the difficulty gradient, so the model learns hard combinations more solidly.

Run the official RL script:

conda activate dyref_rl
cd DyRef/rl
bash scripts/qwen2511-gdpo-rank64-add2k5-csd-siglipv2_flat-sigmoid0.65-focal_loss.sh

The script name is dense with information: rank64 is the LoRA rank, siglipv2 is the vision encoder used for reward, and sigmoid0.65 and focal_loss are implementation details of DAR's difficulty awareness.

Results

After training, validate multi-reference consistency on OmniRef-Bench's test set:

# Run the evaluation inside benchmark/
cd DyRef/benchmark
# The official benchmark script runs 395 samples across 5 reference types × 10 combinations

Success criteria: subject consistency (no face drift), pose alignment, and style fidelity all improve markedly over the base Qwen-Image-Edit-2511, reaching the paper's reported "matches Nano Banana Pro, surpasses Seedream 4.5" level.

FAQ

  • Not enough VRAM: SFT uses LoRA (already the default) and is light on memory; for RL, enable DeepSpeed ZeRO-2/3 sharding. Beginners should start with step 3 inference only.
  • No multi-reference data of your own: reproduce the paper's results directly with the official OmniRef-training set (~14K samples).
  • Want a different base model: the repo also supports FLUX.2-klein-base-9B; see the corresponding script under model_inference/.
  • The training script name is unreadable: that parameter string is actually DyRef's core hyperparameters (rank, sigmoid, focal_loss, etc.); read the RL section of the paper before changing anything.