Alibaba releases Qwen-Audio-3.1 model lineup and cuts AI audio API prices by up to 95%
Alibaba's Qwen team released five new audio models (ASR, ASR-Next, TTS, TTS-Next, and a real-time speech model) with new capabilities (multi-speaker ID with timestamps, emotion/ambient/machine noise detection, cross-language voice transfer, prompt-controlled emotion/style, combined language-model+diffusion generation, low-latency interruptible real-time interaction) and cut API prices: TTS ~70% lower, Realtime ~85% lower, ASR up to 95% lower.
Entities: Alibaba, Qwen, Qwen-Audio-3.1, Qwen Cloud
0 primary
What happened
Alibaba's Qwen team released five new audio models under the Qwen-Audio-3.1 label (ASR, ASR-Next, TTS, TTS-Next, and a real-time speech model), adding features like multi-speaker identification with timestamps, emotion/noise detection, cross-language voice transfer, and low-latency interruptible voice interaction. Alongside the launch, Alibaba cut Qwen Cloud API prices: TTS by roughly 70%, real-time speech by roughly 85%, and ASR by up to 95%. No primary source (official blog post or API pricing page) has been directly verified; the reporting traces back to a secondary write-up citing an X post.
Why it matters
If the price cuts hold up, this is directly useful to anyone buying speech-to-text or text-to-speech APIs at volume: it gives buyers leverage to renegotiate or switch providers, and puts real pricing pressure on OpenAI, ElevenLabs and Deepgram in a market where audio API costs have been falling fast anyway. Developers get a concrete new option with named capabilities (real-time interruptible speech, emotion-aware TTS) worth benchmarking against incumbents. The impact is mostly commercial and immediate rather than strategic; it does not obviously shift who controls the underlying technology.
What is noise
The percentage discounts (70%, 85%, up to 95%) are not anchored to actual per-unit prices, so the real-dollar magnitude cannot be checked, and "up to 95%" likely reflects a best-case tier rather than a typical one. Framing the models as having "empathy" because they detect a low mood is vendor packaging, not a technical claim worth taking at face value. This is also part of a now-routine cadence of audio model releases and price wars, so treat it as incremental rather than a step-change.
Watch next
- 01Find Alibaba's official Qwen Cloud pricing page or blog post and check the actual per-unit rates before and after the cut, not just the percentage claims
- 02Independent benchmarks (e.g., on ASR word-error-rate leaderboards or TTS quality comparisons) showing whether Qwen-Audio-3.1 actually competes with OpenAI Whisper/Realtime, ElevenLabs or Deepgram on accuracy and latency, not just price
- 03Whether OpenAI, ElevenLabs or Deepgram respond with their own price cuts in the next 1-2 months, which would confirm real competitive pressure rather than a one-off promotional move
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680