Ali Qwen-Audio-3.0-TTS Tops Artificial Analysis

Alibaba’s Qwen releases the large speech synthesis model **Qwen-Audio-3.0-TTS**, with the Plus version taking the top spot on the *Artificial Analysis* leaderboard. It supports **16 languages**, **20 dialects**, and **48 kHz studio‑quality audio**, allowing fine‑grained control of emotions and breathing details through **structured tags**.
Alibaba really hasn’t stopped on the speech front this year. On July 15, it had just released the real-time speech dialogue model Qwen-Audio-3.0-Realtime, and today (July 20), Alibaba Qwen has filled in the speech synthesis gap — with Qwen-Audio-3.0-TTS officially launched in Flash and Plus versions. The former focuses on real-time interaction with ~300 ms first‑packet latency, while the latter emphasizes high‑fidelity generation. The Plus version even ranked #1 globally on Artificial Analysis’s TTS leaderboard.
This isn’t one of those demo-and-done models. This time, Alibaba boosted the audio output spec from 24 kHz to 48 kHz. In simpler terms, old TTS systems mostly had “telephone‑like” or “voice assistant” quality — intelligible but obviously synthetic. At 48 kHz, it’s now on par with studio recording and film dubbing standards. Combined with a three‑minute maximum generation length, long‑form uses like podcasts, audiobooks, and dubbing can finally be handled with a single API.

Tag System: Speech Synthesis Starts “Writing Scripts”
The clearest sign of this generation’s TTS philosophy shift is its tag‑control system. Previously, if you wanted a TTS model to laugh, sigh, or sound angry, you either relied on a pile of nested SSML tags or trained a dedicated emotional timbre. Qwen‑Audio‑3.0‑TTS builds this directly into the text layer — write [gasp], [giggles], or [angry] in your content, and the model inserts a gasp, a chuckle, or an angry inflection at that exact spot.
For example, for the same sentence “I can’t believe you actually did it,” you could write:
- “I can’t believe you actually did it.” — neutral statement
- “[gasp] I can’t believe you actually did it.” — gasping in surprise
- “[giggles] I can’t believe you actually did it.” — teasing, amused tone
- “[angry] I can’t believe you actually did it.” — through clenched teeth
Four variants, four completely different emotional directions. For teams making interactive fiction, AI companions, or game dubbing, this level of fine‑grained control matters far more than simply adding new voices. It breaks “emotion” down into atomic operations that can appear anywhere in a sentence — voice synthesis can now be “directed” like a script.
The official release also mentions freestyle instruction compliance, meaning you don’t have to stick to fixed tags: describe naturally (“in a tired, about‑to‑cry voice”), and the model still gets it. This capability already appeared in the Qwen3‑TTS‑Instruct‑Flash series and is now further refined.
Dialects and Minor Languages: Real Investment This Time
Chinese TTS developers have been battling over dialects for the past two years, but most stopped at Cantonese, Sichuanese, or Northeastern Mandarin. Qwen‑Audio‑3.0‑TTS‑Plus supports 20 dialects at once — including Yunnanese, Shaanxi, Shanghainese, and Chongqing varieties — basically the broadest domestic coverage so far.
For multilingual support, it covers 16 major languages, adding Arabic, Vietnamese, Malay, and Filipino. Alibaba’s focus on Southeast Asia and the Middle East markets is obvious — Lazada, AliExpress, and Cainiao all have real‑world needs, so the model isn’t just chasing leaderboard stats.
The accompanying curated voice library will roll out in four categories:
- Instructional‑style voices: optimized for use with the tag system, offering stronger stylistic expression
- 20 dialect voices: directly callable, no secondary tuning required
- Fine‑grained control voices: specialized for breath, laughter, pauses, and other paralinguistics
- 14+ minor‑language voices: aimed at overseas localization
This idea of turning model capabilities into voice resources is smart. Many developers don’t want to tweak parameters or craft prompts — they just want a “northeastern big brother” or “Vietnamese female agent.” The voice library pre‑packages the model’s abilities, lowering the barrier to entry.
Flash vs Plus: Two Parallel Routes
This time, Flash and Plus have clearer roles. Flash achieves ~300 ms first‑packet latency, ideal for real‑time interactions — customer service, companionship, voice agents, etc. Plus emphasizes quality — 48 kHz sampling, long audio, robust acoustics — aimed at content production, podcasts, and dubbing.
A quick note about Qwen‑Audio‑3.0‑Realtime. The real‑time dialogue model launched last week is actually designed to pair with this TTS — Realtime handles “listening + reasoning + responding,” while TTS enhances the sound quality and expressiveness of those responses. For voice‑agent architectures, Realtime + TTS‑Flash forms a complete chain: duplex control, tool calling, emotional feedback, and tag‑level emotion control — covering roughly 90% of dialogue scenarios.

Comparison with GPT‑4o‑TTS and ElevenLabs
You can’t avoid comparing it with OpenAI’s gpt‑4o‑tts and ElevenLabs’ Multilingual v3.
Versus gpt‑4o‑tts: Qwen‑Audio‑3.0‑TTS wins on dialect coverage and tag control. OpenAI’s TTS has impressive expressiveness in English, but almost no Chinese dialect support, nor embedded emotional‑tag control. The Plus version’s 48 kHz quality also surpasses gpt‑4o‑tts’s 24 kHz.
Versus ElevenLabs: ElevenLabs still leads in voice cloning and English expressiveness. Qwen does not emphasize voice cloning this time, opting for a “curated presets + fine‑grained control” strategy instead. In multilingual coverage—especially East and Southeast Asian languages—Qwen’s localization edge is clear. Alibaba Cloud Bailian’s pricing is also typically an order of magnitude cheaper than ElevenLabs, more friendly for high‑concurrency usage.
Artificial Analysis’s top ranking mainly measured naturalness, instruction compliance, and multilingual performance, where Qwen‑Audio‑3.0‑TTS‑Plus indeed stands strong. Still, leaderboards aside, practical A/B testing in your own scenarios matters more.
Application Ideas
With such a model, here’s what you can do:
- Audiobook creation: 3‑minute synthesis, 48 kHz, emotional tags — a 200 k‑word novel can run through the model with human operators simply marking emotional cues. Traditional audiobook studio costs (hundreds to thousands per hour) drop to API call fees.
- Multilingual video dubbing: For languages like Arabic, Vietnamese, or Malay, overseas short‑video teams can generate a draft using the Plus version instead of hiring native actors.
- Game NPC voices: Instead of pre‑recorded lines, use the tag system + voice library; modify scripts and regenerate speech instantly, even dynamically embed player names—once difficult to achieve.
- Dialect customer service: For elderly users in Yunnan, Sichuan, or Shaanxi, dialect voices feel far more natural than Mandarin call agents. Mixed human‑AI setups are now viable.
- Podcasts / media content: A solo podcaster can portray multiple roles using different voices, inserting
[laughs]or[sighs]tags for dramatic effect.
Availability
Qwen‑Audio‑3.0‑TTS is now fully available on Alibaba Cloud Bailian, with both Plus and Flash accessible via API. As part of the open ecosystem, Qwen’s speech models will also be connected to OpenAI Hub, enabling a single key to access GPT, Claude, Gemini, and Qwen‑Audio series models — convenient for teams doing multi‑model benchmarks.
Summary
In the past two years, speech synthesis has evolved from “intelligible” to “human‑like,” and now to “expressive.” The race is no longer about whose timbre sounds most natural, but who offers finer control, broader language coverage, and better audio quality. Qwen‑Audio‑3.0‑TTS advances all three: tag‑level control, 20 dialects & 16 languages, 48 kHz fidelity. That Artificial Analysis #1 placement will likely pressure OpenAI and ElevenLabs in the short term.
Next worth watching: voice cloning and multi‑speaker dialogue generation. An internal Alibaba Cloud Bailian listing already shows a model ID qwen3‑tts‑vc‑2026‑01‑22 for voice cloning — likely not far off.
References
- Ranked first on Artificial Analysis: Alibaba Qwen releases Qwen‑Audio‑3.0‑TTS – IT Home: Details and capabilities of Qwen‑Audio‑3.0‑TTS
- Alibaba launches real‑time speech‑dialogue model Qwen‑Audio‑3.0‑Realtime – IT Home: Companion Realtime model coverage



