Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion

·Toolin Editorial Team

Alibaba's new-generation TTS model controls laughter, gasps, and anger with tags, delivers 48kHz film-grade audio, and tops the global Speech Arena.

Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion

On July 20, 2026, Alibaba's Qwen officially released its next-generation speech synthesis foundation model, Qwen-Audio-3.0-TTS. The problem it set out to solve is no longer "can it speak" but "can it perform" — through structured emotion and tone tags plus natural-language instructions, you can precisely control when to gasp, when to chuckle, and when to sound angry.

Two key results from the official announcement: Qwen-Audio-3.0-TTS-Plus took the global No. 1 spot on the Artificial Analysis Speech Arena with an Elo of 1236, narrowly edging out Simba 3.2 (1234). There is also a Flash version built for real-time interaction, with first-packet latency of about 300ms. If you build voice assistants, audiobooks, podcast voiceovers, or game NPCs, this is currently the most worthwhile tier of Chinese-made TTS to get hands-on with.

Two Versions: Which One to Pick

VersionPositioningLatencyBest For
Qwen-Audio-3.0-TTS-FlashReal-time interactionFirst packet ~300msVoice assistants, customer service, conversational agents
Qwen-Audio-3.0-TTS-PlusHigh-quality generationHigherAudiobooks, film/TV dubbing, content production

Simple rule of thumb: pick Flash for instant responses, pick Plus for beautiful audio.

Core Capabilities: Four Control Methods

1. Structured Emotion/Tone Tags

This is the most visible upgrade in 3.0. Insert square-bracket tags directly into the text, and the model renders the emotional action at the corresponding spot:

She sprang to her feet[gasp], staring at the screen for a long while.
"You actually went and did it..."[giggles] Fine, you win.
I've put up with this for far too long[angry]. Today it all gets said.

Common tags include [gasp], [giggles], and [angry], covering crying, laughing, sighing, pausing, and other emotional actions. It shines for novel narration and story-driven ads.

2. Free-style Natural-Language Instructions

Don't want to memorize tags? Describe the emotion in one sentence and the model performs it on its own:

(Instruction) Read this passage in a tired, self-deprecating tone, pausing slightly between sentences.
(Text) Squeezing into the Monday morning subway, the proposal rejected three times, coffee spilled on a white shirt...

It suits content that needs a certain "vibe" rather than precisely scripted actions, such as short-video scripts and ad voiceovers.

3. 16 Languages + 20 Dialects

It covers 16 major languages including Chinese, English, Japanese, and Korean, plus 20 dialects (such as Cantonese, Sichuanese, and Shanghainese). Cross-language audio content and localized marketing no longer require switching models.

4. 48kHz Film-Grade Audio Quality

A 48kHz sample rate, benchmarked directly against film and TV post-production standards. No more post-hoc noise reduction or upsampling — audio can go straight into the editing pipeline.

How to Integrate: Two Paths via DashScope

Qwen-Audio-3.0-TTS is offered through Alibaba Cloud's Model Studio platform (DashScope), and the same model name supports both WebSocket and HTTP protocols — choose based on your latency requirements.

Path A: Real-Time Speech Synthesis (WebSocket, Low Latency)

Use cases: voice assistants, conversational bots, online customer service.

Call flow (pseudocode):

import dashscope
from dashscope.audio.tts_v2 import SpeechSynthesizer

dashscope.api_key = "sk-xxx"  # Obtain from the Model Studio platform

synthesizer = SpeechSynthesizer(
    model="qwen-audio-3.0-tts-flash",  # Use Flash for real-time
    voice="Cherry",                     # Choose a voice
    callback=your_callback,             # Streaming callback
)
synthesizer.streaming_call("She sprang to her feet[gasp], staring at the screen for a long while.")
synthesizer.streaming_complete()

💡 Tip: The Flash version's first-packet latency is about 300ms, already at conversational levels. If you're building a voice agent, pair it with streaming output from a Qwen LLM for "thinking while speaking."

Path B: Non-Real-Time Speech Synthesis (HTTP, High Quality)

Use cases: batch audiobook synthesis, online education voiceovers, video post-production dubbing.

curl -X POST 'https://dashscope.aliyuncs.com/api/v1/services/audio/tts' \
  -H 'Authorization: Bearer sk-xxx' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "qwen-audio-3.0-tts-plus",
    "input": {
      "text": "She sprang to her feet[gasp], staring at the screen for a long while."
    },
    "parameters": {
      "voice": "Cherry",
      "format": "mp3",
      "sample_rate": 48000
    }
  }'

💡 Tip: For batch content production (such as audiobook versions of an entire book), use HTTP + async tasks to avoid blocking the main thread.

Custom Voices: Voice Cloning

If the default voices aren't enough, the Model Studio platform also offers voice cloning — upload a few seconds to a few tens of seconds of the target voice and train a reusable custom voice. Docs: https://www.alibabacloud.com/help/zh/model-studio/voice-cloning-user-guide.

How It Compares: vs. Doubao and Gege AI

ModelVendorCore StrengthBest For
Qwen-Audio-3.0-TTSAlibabaSpeech synthesis + emotion tags + multilingualDubbing, audiobooks, conversational agents
Doubao Audio Generation 1.0ByteDanceGeneral audio generationSound effects, ambient sound
Gege AIThird partyMusic generationSongwriting, scoring

In short: for TTS go with Qwen-Audio-3.0-TTS, for sound effects go with Doubao, for music go with Gege AI.

Quick-Start Checklist

  1. Register an Alibaba Cloud account and activate Model Studio: https://bailian.console.aliyun.com/
  2. Create a key under "API Key Management"
  3. Copy any of the code samples above, replace sk-xxx, and you'll have your first piece of synthesized speech running
  4. Official sample repository: https://github.com/aliyun/alibabacloud-bailian-speech-demo

FAQ

  • Tags not taking effect, read out literally? Check that you're using a model version that supports tags (Plus / Flash 3.0); the older CosyVoice doesn't recognize structured tags.
  • Latency much higher than 300ms? Confirm you're using the Flash model + WebSocket protocol; HTTP mode cannot guarantee real-time performance.
  • Muffled audio quality? Explicitly set sample_rate to 48000 to avoid dropping to the default 16kHz/24kHz.

Related Links