Back to blogs
Reviews
Comparisons
Benchmarks
Voice AI

Qwen-Audio-3.1 Review: ASR, TTS, Realtime, Price & Is It Worth It? (2026)

September 24, 2026
15 min read
Qwen-Audio-3.1 Review: ASR, TTS, Realtime, Price & Is It Worth It? (2026)
Share:

Qwen-Audio-3.1 Review: Is Qwen's New Audio Stack Ready for Production Voice AI?

Qwen-Audio-3.1 is not one model with one job. It is a new audio stack from Alibaba's Qwen team covering speech recognition, audio understanding, speech synthesis and realtime voice interaction. The September 2026 release brings together upgraded ASR and TTS systems, new Next-generation audio models, and a realtime duplex voice model.

The biggest change is breadth. Qwen-Audio-3.1-ASR is designed for multilingual transcription, speaker diarization, dialect recognition and audio understanding. Qwen-Audio-3.1-TTS-Flash targets low-latency speech synthesis with expressive controls and voice cloning, while Qwen-Audio-3.1-TTS-Next expands into full audio generation with speech, ambience, sound effects and podcast-style content. Realtime-Plus adds duplex speech, function calling, web search and voice cloning.

The pricing has also changed materially. Alibaba's September price-reduction notice cut international Singapore rates for Qwen-Audio-3.1-TTS-Flash to $0.23 per million input tokens and $1.87 per million output tokens. The ASR file-transcription endpoint is currently listed at $0.15 input and $0.47 output per million tokens, while Realtime-Plus uses separate audio and text billing.

Qwen Audio 3.1

QUICK ANSWER

Qwen-Audio-3.1 is best understood as a family rather than a single checkpoint. It combines dedicated ASR, audio-understanding, TTS, full audio generation and realtime speech systems. The ASR stack supports 30 languages and 16 Chinese dialect varieties, while Realtime-Plus supports text and audio input/output, function calling, web search and voice cloning.

The ASR results are the clearest measurable strength. Qwen reports a 4.57 macro error rate on Common Voice 15 across the evaluated languages, 10.38% macro CER on an internal 16-dialect suite, and the highest industry-entity recall in 11 of 15 domains plus one tie. The diarized ASR comparison shows the Next variant leading five of eight listed DER/cpWER metrics.

The TTS side adds instruction-based control, emotion and tone controls, voice cloning, fine-grained expressive tags and long-form synthesis. Qwen’s published TTS evaluation says its 3.1 system ranks first on the cited Artificial Analysis Text-to-Speech leaderboard, while third-party coverage makes clear that the predecessor 3.0 Voice Arena score should not simply be copied onto 3.1.

My verdict: 9.1/10 overall. Qwen-Audio-3.1 is one of the most complete audio model families to test in 2026, especially when a product needs transcription, synthesis and realtime voice from one provider.

1. What Is Qwen-Audio-3.1?

Qwen-Audio-3.1 is Alibaba’s September 2026 refresh of its Qwen audio lineup. Instead of making one monolithic audio model, Qwen separates recognition, synthesis, audio generation and realtime conversation into purpose-built endpoints.

The current family includes Qwen-Audio-3.1-ASR, Qwen-Audio-3.1-ASR-Next, Qwen-Audio-3.1-TTS-Flash, Qwen-Audio-3.1-TTS-Next and Qwen-Audio-3.1-Realtime-Plus. Qwen Cloud places these products in separate ASR, TTS and speech-to-speech categories, which is important when planning an application architecture.

That separation is sensible because speech recognition, speech synthesis and duplex conversation have different latency, context and quality requirements. An ASR endpoint needs strong transcription and diarization. A TTS endpoint needs pronunciation, voice identity and expressive control. A realtime endpoint has to manage turn-taking, audio streaming and tools at the same time.

2. Qwen-Audio-3.1 Family at a Glance

Qwen-Audio-3.1 Family

3. Qwen-Audio-3.1 ASR: The Strongest Measurable Component

The ASR system is the easiest part of the family to evaluate because Qwen publishes a dedicated benchmark and capability site. It is built for multilingual transcription, Chinese dialect recognition, speaker attribution and audio understanding rather than ordinary speech-to-text alone.

Qwen-Audio-3.1-ASR supports 30 languages and 16 Chinese dialect varieties. It can jointly output speaker labels, timestamps and transcript text, and it includes native polishing that removes fillers, repetitions and self-corrections while reorganizing the transcript into cleaner language.

Qwen-Audio-3.1 ASR

4. Qwen-Audio-3.1 ASR Benchmarks

The ASR evaluation covers public and internal tests. Qwen reports a 4.57 macro error rate on Common Voice 15 across the evaluated languages. On its internal 16-dialect evaluation, the model reports a 10.38% macro CER and the lowest CER on 11 of 16 dialect subsets. In industry-domain testing it reports the highest entity recall in 11 of 15 domains plus one tie.

Qwen-Audio-3.1 ASR Benchmarks

Qwen’s comparison shows 81.6 on MMAU, ahead of the listed Seed-2.0-Lite and Gemini-3.5-Flash results, while Gemini-3.1-Pro remains higher on MMSU at 85.2 versus Qwen’s 82.8. That gives the 3.1 ASR model a strong overall audio-understanding profile without suggesting it dominates every audio benchmark.

5. Speaker Diarization: Who Said What?

Speaker diarization is especially valuable for meetings, interviews, support calls and multi-person recordings because a transcription can be word-accurate while still being operationally useless if speaker attribution is wrong.

Qwen-Audio-3.1 uses an end-to-end design that jointly outputs speaker labels, timestamps and text. The published comparison uses AISHELL-4, AliMeeting-test, MLC-SLM and MagicData-RAMC, with the Next variant leading five of the eight displayed DER/cpWER metrics and the SSE variant leading the remaining three.

For production systems, that matters because diarization quality affects the reliability of downstream actions. A meeting assistant that assigns an approval request to the wrong participant can create a more serious problem than a single word transcription error.

6. Native Transcript Polishing

Qwen-Audio-3.1-ASR can also polish the transcript as part of the recognition process. The model removes filler words, repeated fragments and self-corrections, then restructures the wording so the final transcript reads more naturally.

This is useful for content teams because raw ASR and publication-ready text are different outputs. A podcast editor may want a cleaned transcript, while a legal or compliance workflow should preserve the original spoken wording. Keep both forms when traceability matters.

7. Qwen-Audio-3.1 TTS: Speech Generation

Qwen-Audio-3.1-TTS-Flash is positioned as a low-latency speech synthesis model for realtime interaction. The current documentation lists streaming synthesis, free-form instruction following, fine-grained control of emotion, tone, persona, speaking rate and volume, plus voice cloning that is more robust to noisy and reverberant reference audio.

The broader Qwen-Audio-3.1-TTS system adds a production-oriented control layer. Its published technical write-up describes a 12.5 Hz supervised speech tokenizer, an extended training pipeline, 86 fine-grained inline control tags and multilingual/dialect synthesis. It also reports one-pass generation of up to three minutes and up to 48 kHz output through vocoder super-resolution.

Qwen-Audio-3.1 Speech Generation

8. TTS Quality and Benchmark Context

Qwen’s published TTS evaluation says Qwen-Audio-3.1-TTS ranks first on the cited Artificial Analysis Text-to-Speech leaderboard and reports strong results across SEED-TTS-Eval, CV3-Eval, instruction following, long-form generation and acoustic robustness.

There is an important benchmark distinction. Qwen-Audio-3.0-TTS-Plus previously had a measured Voice Arena result around 1,238 Elo, but the 3.1 release is a newer model generation. Current third-party coverage warns against transferring that predecessor score to the new 3.1 endpoint.

So the useful conclusion is that the new TTS system has a strong published evaluation record and an aggressive feature set, but its exact live preference score should always be read from the current 3.1 leaderboard rather than copied from 3.0.

9. Qwen-Audio-3.1 TTS-Next: Beyond Reading Text Aloud

TTS-Next is the most ambitious audio-generation model in the family. Alibaba describes it as an AudioGen model that can generate speech, sound effects, ambience and other complete audio content from text and reference audio. That puts it closer to an audio production model than a conventional TTS endpoint.

Qwen-Audio-3.1 TTS Next

The use cases include podcast generation, cinematic soundscapes, short-form video audio, game sound effects, ambience and multi-speaker dialogue. The output model is therefore useful when the goal is a finished audio scene rather than an isolated spoken sentence.

10. Qwen-Audio-3.1 Realtime Plus

Realtime Plus is the voice-agent endpoint. It accepts text and audio, returns text and audio, and supports function calling, web search and voice cloning. The current context window is 262,144 tokens, with a 245,760-token input maximum and 16,384-token output maximum.

Qwen-Audio-3.1 ASR

One integration detail deserves attention: web search and function calling cannot be enabled together in the same Realtime Plus session. That constraint is documented by Alibaba Cloud and should be part of the architecture if the agent needs both browsing and action tools.

11. Qwen-Audio-3.1 Pricing

Qwen’s audio pricing is now split by workload. Speech recognition, TTS and realtime audio use different billing dimensions, and the international pricing differs from China-region rates.

Qwen-Audio-3.1 Pricing

The ASR and TTS Flash figures reflect current international pricing documentation and Alibaba’s September 22 price reduction. Realtime Plus is billed separately for text and audio. TTS-Next is currently documented with Beijing list pricing, so international teams should check the active regional console before budgeting.

12. Qwen-Audio-3.1 vs Qwen-Audio-3.0

The 3.1 family is not just a quality refresh. Qwen is adding specialized Next models while expanding the existing ASR, TTS and realtime products. That changes the product shape as much as the model quality.

Qwen-Audio-3.1 vs Qwen-Audio-3.0

For a product team, this means 3.1 is better viewed as a platform upgrade. You can select the endpoint that matches the workflow rather than expecting one model to handle every audio problem.

13. Qwen-Audio-3.1 vs Dedicated Voice AI APIs

Qwen’s strongest competitive advantage is breadth. A single vendor can now cover transcription, diarization, speech synthesis, audio generation and realtime speech. That can simplify account management, integration and model routing for voice products.

The tradeoff is specialization. Dedicated providers often focus deeply on one part of the stack. Qwen’s value proposition is therefore the complete audio stack plus competitive pricing, not a claim that every endpoint wins every independent benchmark.

Qwen-Audio-3.1 vs Dedicated Voice AI APIs

14. Voice Cloning and Reference Audio

Voice cloning appears in several parts of the 3.1 family. TTS-Flash supports voice cloning, Realtime Plus supports cloned voices through the voice-cloning API, and TTS-Next accepts up to three reference clips.

Reference quality still matters. Noise, reverberation, multiple speakers and inconsistent recording levels can reduce similarity. The 3.1 TTS documentation specifically highlights better robustness to noisy and reverberant reference audio, which is useful for real-world voice datasets rather than studio-perfect samples.

15. Best Use Cases

Qwen-Audio-3.1 Best Use Cases

16. Limitations You Should Know

  • Qwen-Audio-3.1 is a family of specialized endpoints, not one universal model.
  • Pricing differs by model and deployment region.
  • TTS quality claims and predecessor Voice Arena scores should be kept separate when comparing generations.
  • Realtime web search and function calling cannot be enabled together in one session.
  • Some ASR and TTS Flash endpoints have relatively small token context limits because they are dedicated speech services.
  • The 3.1 family is hosted through Qwen Cloud/Alibaba Cloud rather than exposed as a downloadable open-weight audio stack.
  • Raw transcription and polished transcription serve different audit requirements and should not automatically be merged.

The cleanest architecture is to select the endpoint by the job rather than forcing one Qwen-Audio model to do everything.

  • Use ASR-Flash-Filetrans for uploaded meetings, calls and long-form recordings.
  • Use the realtime ASR path when partial transcription is required during a live interaction.
  • Use TTS-Flash for low-latency assistant speech and expressive voice control.
  • Use TTS-Next when the output is a complete podcast, scene or soundscape rather than simple speech.
  • Use Realtime Plus when the application requires two-way spoken conversation with tools or web search.
  • Keep raw ASR output separate from cleaned/polished transcripts when traceability matters.
  • Measure cost per finished minute or completed voice interaction, not only token price.

For the broader agent architecture, see What Is an AI Agent?.

18. How to Evaluate Qwen-Audio-3.1 Yourself

A voice stack should be evaluated with production-like recordings, not only public benchmark numbers. Build small test sets for each workload you actually plan to ship.

Evaluate Qwen-Audio-3.1 Yourself

A meeting assistant may care more about speaker attribution than raw WER. A customer-service agent may care more about interruption latency. A podcast tool may care most about long-form voice consistency. The best model choice is therefore the endpoint that wins your actual task-level evaluation.

19. Is Qwen-Audio-3.1 Worth It?

Yes, particularly if you want a broad speech stack from one provider. The family now spans transcription, diarization, audio understanding, expressive TTS, complete audio generation and realtime voice agents, while selected endpoint prices have fallen sharply.

The strongest evidence is on the ASR side. The 30-language and 16-dialect coverage, low reported error rates, speaker attribution and audio-understanding results make Qwen-Audio-3.1-ASR a compelling option for production transcription systems.

TTS-Next is the most differentiated addition because it moves beyond speech into complete audio creation. Realtime Plus is equally important for developers building voice agents, since it combines duplex speech with tools, web search and voice cloning.

The main architectural cost is that you need to select among specialized APIs. That is manageable for a production engineering team, but it means Qwen-Audio-3.1 is better thought of as an audio platform than a single drop-in model.

20. Final Verdict

Qwen-Audio-3.1 is one of the most complete audio releases of September 2026 because it treats voice as a stack rather than a single model. The family covers ASR, speaker diarization, audio understanding, expressive TTS, soundscape generation and realtime speech-to-speech interaction.

The ASR numbers are particularly convincing. Qwen reports 30 languages, 16 Chinese dialect varieties, a 4.57 macro error rate on Common Voice 15, 10.38% macro CER across its internal dialect suite and the highest industry entity recall in 11 of 15 domains plus one tie.

The TTS side brings strong controllability and a broader creative target. Qwen’s published technical evaluation reports a first-place position on the cited Artificial Analysis TTS leaderboard, together with fine-grained controls, voice cloning, long-form synthesis and robust reference-audio handling.

Realtime Plus makes the release more relevant to AI agents. With 262K context, duplex speech, function calling, web search and voice cloning, it can act as the voice layer inside a tool-using assistant rather than just a speech endpoint.

The current international pricing is also attractive for high-volume applications. ASR-Flash-Filetrans is listed at $0.15 input and $0.47 output per million tokens, TTS-Flash at $0.23 input and $1.87 output, and Realtime Plus at $0.80 text input, $6.40 audio input, $6.40 text output and $24 audio output per million tokens in Singapore.

My rating: 9.4/10 for ASR, 9.1/10 for TTS, 9.2/10 for realtime agents, 9.3/10 for price-to-capability and 9.1/10 overall.

Bottom line: Qwen-Audio-3.1 is worth testing if you are building speech products in 2026. Its biggest advantage is not that every endpoint wins every benchmark. It is that Qwen now provides a broad, production-oriented audio stack that can cover the complete loop from hearing to understanding to speaking.

Frequently Asked Questions

What is Qwen-Audio-3.1?

Qwen-Audio-3.1 is Alibaba’s September 2026 audio family covering ASR, audio understanding, TTS, full audio generation and realtime voice.

How many models are in Qwen-Audio-3.1?

The current family includes ASR, ASR-Next, TTS-Flash, TTS-Next and Realtime-Plus.

How many languages does Qwen-Audio-3.1 ASR support?

The official ASR evaluation site lists 30 languages and 16 Chinese dialect varieties.

Does Qwen-Audio-3.1 support speaker diarization?

Yes. The ASR family jointly outputs speaker labels, timestamps and transcription text.

What is the Qwen-Audio-3.1 ASR benchmark?

Qwen reports 4.57 macro error rate on Common Voice 15 and 10.38% macro CER on its internal 16-dialect evaluation.

What is Qwen-Audio-3.1 TTS?

It is the speech-synthesis part of the family, with instruction-based control, expressive tags, voice cloning and long-form capabilities.

Does Qwen-Audio-3.1 support voice cloning?

Yes. The TTS and Realtime products support voice cloning or reference-based voice control.

What is Qwen-Audio-3.1 TTS-Next?

It is the AudioGen-oriented endpoint that can generate speech plus ambience, sound effects and other complete audio content.

What is Qwen-Audio-3.1 Realtime Plus?

It is a duplex speech model with text and audio input/output, 262K context, function calling, web search and voice cloning.

How much does Qwen-Audio-3.1 cost?

International pricing varies by endpoint. ASR-Flash-Filetrans is $0.15/$0.47 per million input/output tokens, TTS-Flash is $0.23/$1.87, and Realtime Plus is separately priced for text and audio.

Is Qwen-Audio-3.1 open source?

The current 3.1 hosted family is provided through Qwen Cloud and Alibaba Cloud Model Studio rather than as a downloadable open-weight audio stack.

Is Qwen-Audio-3.1 worth it?

Yes, especially for production transcription, speech synthesis, audio generation and realtime voice agents that benefit from one provider covering several audio tasks.

Resources & Community

Join our community of 70,000+ AI enthusiasts and learn to build powerful AI applications. Whether you are a beginner or an experienced developer, Build Fast with AI helps you understand and implement AI in your projects.

Agentic AI Launchpad 2026

A structured 6-week cohort program that takes you from AI basics to building and deploying real-world agentic AI systems. Includes live sessions, expert mentorship, project reviews and a builder community network.

Ready to go from learning to building? Join the next cohort: Agentic AI Launchpad 2026

Free AI Resources

Access free tools, workshops and micro-learning to keep building.

References

Share: