Alibaba Tongyi Lab Releases Qwen-Audio-3.0-TTS With 16-Language Support Across Flash and Plus Tiers

Alibaba's Tongyi Lab has released Qwen-Audio-3.0-TTS, a hosted text-to-speech model available via API in Flash (faster, lower cost) and Plus (higher quality) tiers, covering 16 languages. This positions Qwen-Audio-3.0-TTS as a direct competitor to ElevenLabs, OpenAI TTS, and Google's TTS offerings, with a notably broad language coverage that could be advantageous for multilingual product teams. The tiered API structure makes it accessible for both prototyping and production workloads, and Alibaba's competitive pricing history suggests this will undercut Western alternatives. For developers building voice interfaces, accessibility tools, or multilingual products, this is worth benchmarking immediately—especially if coverage of Asian languages has been a pain point. The model's availability as a hosted API means no self-hosting overhead, which lowers the integration cost substantially.
Read original source ↗Part of the 2026-07-21 digest→