Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion
Alibaba's new-generation TTS model controls laughter, gasps, and anger with tags, delivers 48kHz film-grade audio, and tops the global Speech Arena.


Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion
Alibaba's new-generation TTS model controls laughter, gasps, and anger with tags, delivers 48kHz film-grade audio, and tops the global Speech Arena.
On July 20, 2026, Alibaba's Qwen officially released its next-generation speech synthesis foundation model, Qwen-Audio-3.0-TTS. The problem it set out to solve is no longer "can it speak" but "can it perform" — through structured emotion and tone tags plus natural-language instructions, you can precisely control when to gasp, when to chuckle, and when to sound angry.
Two key results from the official announcement: Qwen-Audio-3.0-TTS-Plus took the global No. 1 spot on the Artificial Analysis Speech Arena with an Elo of 1236, narrowly edging out Simba 3.2 (1234). There is also a Flash version built for real-time interaction, with first-packet latency of about 300ms. If you build voice assistants, audiobooks, podcast voiceovers, or game NPCs, this is currently the most worthwhile tier of Chinese-made TTS to get hands-on with.
Two Versions: Which One to Pick
| Version | Positioning | Latency | Best For |
|---|---|---|---|
| Qwen-Audio-3.0-TTS-Flash | Real-time interaction | First packet ~300ms | Voice assistants, customer service, conversational agents |
| Qwen-Audio-3.0-TTS-Plus | High-quality generation | Higher | Audiobooks, film/TV dubbing, content production |
Simple rule of thumb: pick Flash for instant responses, pick Plus for beautiful audio.
Core Capabilities: Four Control Methods
1. Structured Emotion/Tone Tags
This is the most visible upgrade in 3.0. Insert square-bracket tags directly into the text, and the model renders the emotional action at the corresponding spot:
She sprang to her feet[gasp], staring at the screen for a long while.
"You actually went and did it..."[giggles] Fine, you win.
I've put up with this for far too long[angry]. Today it all gets said.Common tags include [gasp], [giggles], and [angry], covering crying, laughing, sighing, pausing, and other emotional actions. It shines for novel narration and story-driven ads.
2. Free-style Natural-Language Instructions
Don't want to memorize tags? Describe the emotion in one sentence and the model performs it on its own:
(Instruction) Read this passage in a tired, self-deprecating tone, pausing slightly between sentences.
(Text) Squeezing into the Monday morning subway, the proposal rejected three times, coffee spilled on a white shirt...It suits content that needs a certain "vibe" rather than precisely scripted actions, such as short-video scripts and ad voiceovers.
3. 16 Languages + 20 Dialects
It covers 16 major languages including Chinese, English, Japanese, and Korean, plus 20 dialects (such as Cantonese, Sichuanese, and Shanghainese). Cross-language audio content and localized marketing no longer require switching models.
4. 48kHz Film-Grade Audio Quality
A 48kHz sample rate, benchmarked directly against film and TV post-production standards. No more post-hoc noise reduction or upsampling — audio can go straight into the editing pipeline.
How to Integrate: Two Paths via DashScope
Qwen-Audio-3.0-TTS is offered through Alibaba Cloud's Model Studio platform (DashScope), and the same model name supports both WebSocket and HTTP protocols — choose based on your latency requirements.
Path A: Real-Time Speech Synthesis (WebSocket, Low Latency)
Use cases: voice assistants, conversational bots, online customer service.
- Protocol: WebSocket streaming input/output
- Endpoint:
wss://dashscope.aliyuncs.com/api-ws/v1/inference/ - Docs: https://help.aliyun.com/zh/model-studio/realtime-tts-user-guide
Call flow (pseudocode):
import dashscope
from dashscope.audio.tts_v2 import SpeechSynthesizer
dashscope.api_key = "sk-xxx" # Obtain from the Model Studio platform
synthesizer = SpeechSynthesizer(
model="qwen-audio-3.0-tts-flash", # Use Flash for real-time
voice="Cherry", # Choose a voice
callback=your_callback, # Streaming callback
)
synthesizer.streaming_call("She sprang to her feet[gasp], staring at the screen for a long while.")
synthesizer.streaming_complete()💡 Tip: The Flash version's first-packet latency is about 300ms, already at conversational levels. If you're building a voice agent, pair it with streaming output from a Qwen LLM for "thinking while speaking."
Path B: Non-Real-Time Speech Synthesis (HTTP, High Quality)
Use cases: batch audiobook synthesis, online education voiceovers, video post-production dubbing.
- Protocol: HTTP POST
- Endpoint:
https://dashscope.aliyuncs.com/api/v1/services/audio/tts - Model:
qwen-audio-3.0-tts-plus - Docs: https://help.aliyun.com/zh/model-studio/non-realtime-tts-user-guide
curl -X POST 'https://dashscope.aliyuncs.com/api/v1/services/audio/tts' \
-H 'Authorization: Bearer sk-xxx' \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen-audio-3.0-tts-plus",
"input": {
"text": "She sprang to her feet[gasp], staring at the screen for a long while."
},
"parameters": {
"voice": "Cherry",
"format": "mp3",
"sample_rate": 48000
}
}'💡 Tip: For batch content production (such as audiobook versions of an entire book), use HTTP + async tasks to avoid blocking the main thread.
Custom Voices: Voice Cloning
If the default voices aren't enough, the Model Studio platform also offers voice cloning — upload a few seconds to a few tens of seconds of the target voice and train a reusable custom voice. Docs: https://www.alibabacloud.com/help/zh/model-studio/voice-cloning-user-guide.
How It Compares: vs. Doubao and Gege AI
| Model | Vendor | Core Strength | Best For |
|---|---|---|---|
| Qwen-Audio-3.0-TTS | Alibaba | Speech synthesis + emotion tags + multilingual | Dubbing, audiobooks, conversational agents |
| Doubao Audio Generation 1.0 | ByteDance | General audio generation | Sound effects, ambient sound |
| Gege AI | Third party | Music generation | Songwriting, scoring |
In short: for TTS go with Qwen-Audio-3.0-TTS, for sound effects go with Doubao, for music go with Gege AI.
Quick-Start Checklist
- Register an Alibaba Cloud account and activate Model Studio: https://bailian.console.aliyun.com/
- Create a key under "API Key Management"
- Copy any of the code samples above, replace
sk-xxx, and you'll have your first piece of synthesized speech running - Official sample repository: https://github.com/aliyun/alibabacloud-bailian-speech-demo
FAQ
- Tags not taking effect, read out literally? Check that you're using a model version that supports tags (Plus / Flash 3.0); the older CosyVoice doesn't recognize structured tags.
- Latency much higher than 300ms? Confirm you're using the Flash model + WebSocket protocol; HTTP mode cannot guarantee real-time performance.
- Muffled audio quality? Explicitly set
sample_rateto48000to avoid dropping to the default 16kHz/24kHz.
Related Links
- Real-time speech synthesis docs: https://help.aliyun.com/zh/model-studio/realtime-tts-user-guide
- Non-real-time speech synthesis docs: https://help.aliyun.com/zh/model-studio/non-realtime-tts-user-guide
- Python SDK reference: https://help.aliyun.com/zh/model-studio/cosyvoice-tts-python-sdk
- Official sample repository: https://github.com/aliyun/alibabacloud-bailian-speech-demo