VeraRetouch: A 0.6B-Parameter AI Retouching Model and a New Path to On-Device Mobile Deployment
vivo and Zhejiang University release VeraRetouch, a lightweight retouching framework built on a 0.6B vision-language model that supports auto, style, and param retouching, processing on an iPhone in about 13 seconds


VeraRetouch: A 0.6B-Parameter AI Retouching Model and a New Path to On-Device Mobile Deployment
vivo and Zhejiang University release VeraRetouch, a lightweight retouching framework built on a 0.6B vision-language model that supports auto, style, and param retouching, processing on an iPhone in about 13 seconds
VeraRetouch is a lightweight AI retouching framework released by vivo BlueImage Lab together with Zhejiang University and Zhejiang Lab. It uses a 0.6B-parameter vision-language model (VLM) as the "retouching brain," paired with a fully differentiable Retouch Renderer as the "retouching executor," to deliver professional-grade tone and color optimization entirely on a phone. Total parameters come to only about 0.63B, far smaller than mainstream solutions like Flux.1 Kontext.
The Core Problem: Why a New Retouching Framework
Existing AI retouching solutions have several key pain points:
- Professional retouching tools are complicated to operate, while one-click filters feel stylistically heavy-handed
- Reasoning-based retouching drags in models that are too heavy, unsuited to mobile deployment
- External retouching software is not differentiable, so model parameters cannot be optimized directly through end-to-end training
VeraRetouch's key innovation: instead of treating professional retouching tools as an external black box, it replaces traditional color-grading and lighting operations with a fully differentiable Retouch Renderer.

VeraRetouch uses a 0.6B VLM as the "retouching brain" and the Retouch Renderer as the "retouching executor."
Three Retouching Modes
VeraRetouch defines three classes of retouching tasks, covering the full range of needs from "one-click enhancement" to "precise control":
Auto-Retouch
The user only needs to provide a photo; the model automatically analyzes the lighting and color problems in the frame and generates a retouching plan. The goal is not to slap on a filter, but to improve the overall look while preserving the original content.
Style-Retouch
The user can describe the desired style in natural language, such as "warm autumn feel," "cool Japanese-style translucency," or "dark moody film look." The model combines image content with textual intent to reason out a concrete color-grading direction.
Param-Retouch
The model retouches according to explicit parameter instructions, such as contrast, exposure, color temperature, or saturation. The emphasis is on the controllability and reproducibility of the adjustments.
Technical Architecture
A Three-Dimensional Breakdown of the Retouching Space
The research team breaks the retouching space into three relatively independent control dimensions:
- Lighting: exposure-, shadow-, and highlight-related adjustments
- Global Color: global color adjustments such as color temperature, tint, and overall color tendency
- Specific Color: fine adjustments targeting specific color channels such as red, orange, and blue
This breakdown closely mirrors professional retouching workflows, making the model's output more interpretable and more stable.
End-to-End Differentiable
The model's final-layer hidden state is fed into an MLP Retouch Adaptor, which aligns it to a continuous control latent the Retouch Renderer understands. The entire retouching process happens inside the model, supporting end-to-end pixel-level training.
The DAPO-AE Post-Training Strategy
Through format rewards, image-similarity rewards, and aesthetic rewards, the model is guided to produce more natural retouching results while maintaining instruction consistency.
The Dataset: AetherRetouch-1M+
To solve the scarcity of professional retouching data, the research team built a million-scale, multi-task professional retouching dataset:
- Auto retouching: takes a "reverse degradation" approach, starting from high-quality photos and generating "unretouched" versions in reverse
- Style retouching: curated 5030 online style presets covering 11 major categories and 193 fine-grained subcategories
- Param retouching: training data generated by randomly sampling parameter combinations around three classes of operations
Benchmark Results
Quality Metrics
| Task | Metric | VeraRetouch | vs. Flux.1 Kontext |
|---|---|---|---|
| Auto-Retouch (FiveK-Bench) | PSNR | 26.85 dB | +1.08 dB |
| Param-Retouch | PSNR | 30.18 dB | clearly ahead |
| Style-Retouch (Aether-Bench) | PSNR/SSIM/LPIPS | best on multiple metrics | leading |
Speed Comparison
| Platform | VeraRetouch | Flux.1 Kontext | JarvisArt |
|---|---|---|---|
| H20 GPU (512p) | 6.90s | 16.78s | 14.31s |
| MacBook Air M4 | about 7.46s | - | - |
| iPhone 16 Pro | about 13.56s | - | - |

On FiveK-Bench, VeraRetouch-DAPO-AE reaches 26.85 dB PSNR, leading on multiple metrics.

On style-retouching tasks, VeraRetouch achieves the best results on multiple metrics.
User Study
A blind evaluation by 38 participants gave VeraRetouch the highest ratings in visual aesthetics, instruction consistency, and texture preservation. The DAPO-AE post-training delivered a 61.62% preference rate.
Practical Information
- Paper: https://arxiv.org/pdf/2604.27375
- Project page: https://apollo-yi.github.io/VeraRetouch/
- Code: https://github.com/OpenVeraTeam/VeraRetouch
- Total parameters: about 0.63B
- Target platforms: on-device mobile deployment (iPhone, MacBook, and similar)
Who It's For
- Phone makers: integrating on-device AI retouching capabilities
- Photo app developers: a lightweight retouching SDK
- Individual developers: locally deployed AI retouching with no cloud calls