On September 23, Alibaba's Qwen team announced Qwen-Audio-3.1, a new series of voice AI models. It updates speech recognition, text-to-speech and real-time conversation, and adds ASR-Next, which understands speech and ambient sound, and TTS-Next, which generates voices and sound effects together, for a lineup of five models.
Prices for voice APIs are also falling, giving developers more room to choose based on use case and budget. But a look at how Qwen compares with competitors shows the verdict depends on which part of a conversation Qwen is strong in, and on which model a given low price actually applies to.
To judge whether it can replace GPT or Gemini, you need to look separately at recognition accuracy, behavior during conversation, and the quality of the generated speech.
From speech recognition to speech generation: the roles of the five models
Qwen-Audio-3.1 includes a model that converts speech to text, a model that responds by voice, and a model that generates audio content including dialogue and sound effects. Even though all are called "voice AI," their inputs, outputs and intended uses differ considerably.
| Model line | Main role | Key features in this release | Notes for Japanese use |
|---|---|---|---|
| ASR | Converts speech to text | Dialect recognition, handling of technical terms, speaker diarization in the non-streaming version | Japanese is among the recognized languages |
| ASR-Next | Understands speech and surrounding sounds | Speaker diarization, descriptions of environmental sounds, identifying when a sound occurs, answering questions about audio | Availability terms for a dedicated API need to be checked separately |
| TTS | Turns text into speech | Emotion and speaking style set by natural-language instructions, voice reproduction across languages | The official demo includes Japanese |
| TTS-Next | Generates voices, sound effects and ambient sound together | Multi-speaker conversations, podcasts, scene-appropriate audio generation | The public API supports Chinese and English |
| Realtime | Interacts by voice in real time | Full-duplex conversation that keeps listening while the other person speaks, tool calling, voice cloning | The user guide includes Japanese |
Based on the official ASR page, the official TTS demo, the TTS-Next specifications and the Realtime user guide. The features the series as a whole advertises are not the same as what each individual API actually offers in practice, including supported languages.
For meeting transcription, for example, speaker diarization, which distinguishes who spoke when, is useful. For video and podcast production, TTS-Next, which can generate not only dialogue but also footsteps and room reverberation in one pass, is a candidate.
Just because the standard TTS has a Japanese demo does not mean TTS-Next can generate Japanese audio content.
Understanding conversation and reacting to interruptions are separate skills
In Alibaba's evaluations, Qwen-Audio-3.1-Realtime outperforms GPT-Realtime-2 on many items. However, results differ between tests where the model answers spoken input in text and tests where it answers spoken input in speech. The Qwen team's technical report evaluates these two types separately.
| Evaluation item | GPT-Realtime-2 | SeedDuplex 1.2.6.1 | Qwen 3.0 Realtime | Qwen 3.1 Realtime |
|---|---|---|---|---|
| Following instructions over multiple spoken turns: Audio MultiChallenge | 50.33% | 41.91% | 47.12% | 52.21% |
| Spoken reasoning problems: Big Bench Audio | 93.30% | 80.50% | 98.80% | 98.50% |
| Task completion: in-house half-duplex S2T evaluation using τ-Voice | 43.8% | 40.3% | 78.4% | 82.0% |
| Function calling from speech: SpeechFCEval | 70.74% | 54.73% | 83.28% | 86.00% |
| Spoken-response evaluation: EVA-A Pass | 50.40% | Not listed | 43.10% | 47.74% |
Higher is better for all. Source: Tables 2 and 6 of the report. The official names of the Qwen models are Qwen-Audio-3.0-Realtime and Qwen-Audio-3.1-Realtime, abbreviated in the table. GPT-Realtime-2 was generally run at low-effort settings. The top four rows are speech-in, text-out (S2T); EVA-A is a speech-in, speech-out (S2S) evaluation.
Audio MultiChallenge was scored with GPT-4o-mini rather than the official o4-mini, so the results cannot be compared directly with scores from the official grader. τ-Voice also differs from the official full-duplex S2S method.
Qwen 3.1 improved on tests of following instructions added mid-conversation and of turning spoken requests into tool operations. On EVA-A Pass, however, GPT-Realtime-2 comes out ahead. On Big Bench Audio, it also fell slightly below the older Qwen 3.0.
So it is not possible to conclude from this table alone that Qwen is ahead in every aspect of voice conversation.
The training approach also shows an emphasis on actually completing tasks, not merely understanding instructions.
Alibaba runs the model in environments equipped with tools and databases, checks whether permitted changes were actually completed, and uses those results in training. Generating a correctly formatted tool call does not by itself mean the job was done.
It also combines training on understanding linguistic instructions with training on speech-specific matters such as emotional expression and conversational timing.
Still, an AI that converses naturally with people must also avoid replying to speech not directed at it, and must stop talking immediately when the user interrupts. Here a different performance gap appears.
| Full-duplex conversation metric | GPT-Realtime-2 | SeedDuplex 1.2.6.1 | Qwen 3.0 | Qwen 3.1 |
|---|---|---|---|---|
| Rate of wrongly reacting to surrounding speech | 72.0% | 61.0% | 73.0% | 13.0% |
| Rate of wrongly reacting when the user talks to someone else | 60.0% | 81.0% | 13.0% | 3.0% |
| Time to respond after the user interrupts | 1.872 s | 2.108 s | 1.9432 s | 1.954 s |
Table 10 of the report, S2S evaluation using Full-Duplex-Bench v1.5. Lower is better for all. Percentages are the original ratios multiplied by 100. The top two rows show how often the AI reacted to speech it should not have responded to.
{
"type": "bar",
"title": "Time from user interruption until the AI stops speaking",
"unit": "seconds, lower is better",
"categories": ["GPT-Realtime-2", "SeedDuplex 1.2.6.1", "Qwen-Audio-3.0-Realtime", "Qwen-Audio-3.1-Realtime"],
"series": [
{"name": "Time to stop speaking", "values": [0.383, 1.419, 1.041, 1.116]}
],
"caption": "Alibaba's Full-Duplex-Bench v1.5 evaluation. Speech in, speech out. GPT-Realtime-2 at low-effort settings. Not the latency of the overall real-world service.",
"source": "Qwen-Audio-3.1-Realtime technical report, Table 10"
}Qwen 3.1 improved greatly at not mistaking surrounding conversation for speech aimed at it. On the other hand, the time from a user's interruption until the AI stops speaking is longer than for GPT-Realtime-2.
In the same Alibaba evaluation, Qwen's false-reaction rate to surrounding speech was 59 percentage points lower than GPT-Realtime-2's, but the time to stop speaking when a user interrupted was 0.733 seconds longer. The differences are "72.0 − 13.0" and "1.116 − 0.383" respectively.
The ability to tell whether speech is directed at it and the ability to react quickly to a user's interruption need to be evaluated separately.
This comparison comes from test results in an Alibaba preprint; it has not been reproduced by third parties, nor has statistical significance been confirmed. Even so, it shows that the metric to prioritize changes depending on whether you want to reduce false reactions to surrounding conversation at a support desk, or want the AI to stop talking the instant a user cuts in.
Chinese dialect recognition improves; keep it separate from Japanese accuracy
The speech recognition model Qwen-Audio-3.1-ASR supports 30 languages and 16 Chinese dialects, according to the official description. In competitor comparisons, Alibaba has published concrete results for Chinese dialects.
Turning the comparison chart on public datasets into a table shows items where Tencent comes out ahead.
| Public dataset / subset | Doubao-ASR | Tencent Hy-ASR-3.0-preview | Qwen-Audio-3.1-ASR |
|---|---|---|---|
| KeSpeech · Beijing | 5.13% | 4.25% | 3.62% |
| KeSpeech · Ji-Lu | 5.87% | 4.11% | 3.98% |
| KeSpeech · Jiang-Huai | 8.24% | 6.39% | 6.33% |
| KeSpeech · Jiao-Liao | 6.37% | 4.31% | 4.20% |
| KeSpeech · Lan-Yin | 6.23% | 4.44% | 4.49% |
| KeSpeech · Northeastern | 7.84% | 4.69% | 4.79% |
| KeSpeech · Southwestern | 4.95% | 3.50% | 3.63% |
| KeSpeech · Zhongyuan | 4.90% | 3.02% | 3.03% |
| KeSpeech · Mandarin | 2.90% | 1.89% | 1.78% |
| WSYue · long audio | 16.78% | 8.90% | 8.77% |
| WSYue · short audio | 9.47% | 5.24% | 5.42% |
The metric is character error rate (CER); lower is better. Values are taken from the same chart Alibaba presented. Note, however, that KeSpeech is an AST evaluation that converts dialect speech into standard Chinese, while WSYue is an ASR evaluation that transcribes Cantonese speech, so they do not measure the same process.
Qwen recorded the lowest error rate on 6 of 11 items, and Tencent on 5.
Simply counting "items won," however, treats the tiny margin on Zhongyuan and the margin on long Cantonese audio as the same single item. Nor can results that combine translation and transcription be read as meaning it will recognize any speech with this accuracy.
Alibaba's evaluation of 16 Chinese dialects using in-house data shows an even larger gap.
データを表で見る
| Simple average CER across 16 dialects (%, lower is better) | |
|---|---|
| Doubao-ASR | 20.23 |
| Tencent Hy-ASR-3.0-preview | 17.13 |
| Qwen-Audio-3.1-ASR | 10.38 |
Qwen's average CER was 10.38%, lower than both of the other models compared. However, since this is a result from Alibaba's in-house speech data, the same gap may not appear in real-world environments.
For Japanese, the Realtime technical report lists Japanese accuracy on the multilingual Big Bench Audio as 89.3% for Qwen 3.1 and 79.8% for GPT-Realtime-2.
This, though, measures whether the model can correctly answer reasoning problems posed by voice. It is not a figure for how accurately Japanese proper nouns are transcribed, or how natural the pronunciation of read-aloud Japanese is.
The specification "supports Japanese" and the accuracy you actually get with Japanese need to be checked separately.
Third-party TTS ratings are for the old version and don't apply to 3.1
Qwen's official TTS demo showcases changing emotion and speaking style through natural-language instructions and preserving voice characteristics when switching languages.
On the technical side, it describes a mechanism that converts audio into 12.5 tokens per second and a method of training the language-model part and the speech-generation part in stages. The aim is to keep the number of generated audio tokens down while balancing accuracy of content with expressive voice quality.
For sound-quality comparisons with competitors, the independent evaluator Artificial Analysis's Provider Voice Arena is a useful reference. However, the Qwen model listed as of September 24 is Qwen-Audio-3.0-TTS-Plus, not 3.1.
| Model | Elo | 95% CI width | Comparison samples | Listed price per 1M characters |
|---|---|---|---|---|
| Cartesia Sonic 3.6 | 1273 | ±17 | 1,757 | $49.0 |
| Google Gemini 3.8 Flash TTS | 1260 | ±17 | 1,999 | $16.5 |
| Qwen-Audio-3.0-TTS-Plus (old version) | 1259 | ±17 | 1,447 | $27.6 |
| Inworld Realtime TTS-2 | 1245 | ±18 | 1,239 | $20.8 |
| SpeechifyAI Simba 3.2 | 1237 | ±14 | 2,391 | $6.6 |
| ElevenLabs v3 Conversational | 1196 | ±15 | 1,844 | $50.0 |
| MiniMax Speech 2.8 HD | 1168 | ±11 | 4,414 | $100.0 |
| ElevenLabs Eleven v3 | 1167 | ±11 | 4,345 | $100.0 |
As of September 24, 2026. Accent and Category are both set to All. Each company's default voice reads the same text, and listeners who don't know the model names choose the voice that sounds more natural.
Elo is a measure of how often a voice tends to be chosen, not a percentage of correct answers. The prices are what the organization lists as the cost of generating 1 million characters at each company's API default settings, in a different unit from the per-million-token prices discussed below.
Although the old Qwen ranks near the top, the confidence intervals of the top three models overlap. The displayed ranking alone does not show a clear performance difference.
Nor can the naturalness of Japanese be judged from this evaluation, which covers US and UK accents. Moreover, the paper linked from the 3.1 demo page is titled 3.0-TTS. The old version's strong results should not be treated as measured results for the new 3.1.
TTS-Next differs in purpose from ordinary text reading.
Because it can generate not only dialogue but also sound effects and ambient sound together, it may reduce the work of editing and layering multiple audio sources afterward.
The public API lets you specify up to three reference audio clips, but supported languages are Chinese and English, and it is non-streaming, returning a finished audio file. The generated audio limit is 240 seconds for podcasts and 120 seconds for everything else.
The right model differs between reading Japanese aloud in real time and producing audio content that includes conversation and sound effects.
Which model is really cheaper? Comparing prices with GPT and Gemini
Qwen-Audio-3.1-Realtime-Plus's audio pricing is lower than GPT-Realtime-2.1's but higher than Gemini 3.8 Live's. Looking only at the size of the price cuts makes this relationship easy to miss.
| Conversation API | Text input | Audio input | Text output | Audio output |
|---|---|---|---|---|
| Qwen-Audio-3.1-Realtime-Plus | $0.80 | $6.40 | $6.40 | $24.00 |
| Qwen-Audio-3.0-Realtime-Flash (old version) | $0.23 | $0.93 | $0.70 | $1.87* |
| OpenAI GPT-Realtime-2.1 | $4.00 | $32.00 | $24.00 | $64.00 |
| OpenAI GPT-Realtime-2.1-mini | $0.60 | $10.00 | $2.40 | $20.00 |
| Google Gemini 3.8 Live | $0.75 | $3.00 | $4.50 | $12.00 |
| Google Gemini 3.1 Flash Live Preview | $0.75 | $3.00 | $4.50 | $12.00 |
All columns are per 1 million tokens. Based on the Alibaba Cloud Plus pricing, the Flash price-revision notice, the OpenAI pricing page and the Google pricing page as of September 24, 2026.
Qwen uses the international Singapore pricing, and no company's cache discounts or external tool fees are included. *Flash's $1.87 is the tier described as "text + audio output" in the revision notice.
{
"type": "bar",
"title": "Audio unit prices for real-time conversation APIs",
"unit": "USD per 1M tokens",
"categories": ["Qwen 3.1 Realtime Plus", "Qwen 3.0 Realtime Flash (old version)", "GPT-Realtime-2.1", "GPT-Realtime-2.1-mini", "Gemini 3.8 Live"],
"series": [
{"name": "Audio input", "values": [6.4, 0.93, 32, 10, 3]},
{"name": "Audio output", "values": [24, 1.87, 64, 20, 12]}
],
"caption": "September 24, 2026. Qwen at Singapore pricing. Flash output is the text + audio tier. Audio token counts do not mean the same duration across companies, so this is not a cost comparison per minute of conversation.",
"source": "Official pricing pages and revision notices from Alibaba Cloud, OpenAI and Google"
}Qwen 3.1 Plus has lower audio input and output prices than GPT-Realtime-2.1, but higher than Google's. Compared with OpenAI's mini, Qwen is cheaper for audio input, while mini is cheaper for audio output.
Actual costs vary with the ratio of time the user speaks to time the AI responds.
One thing to be careful about is not confusing the models used for the performance comparison with those used for the price comparison.
Alibaba's evaluations in the first half targeted GPT-Realtime-2, while the OpenAI model in the current price list is GPT-Realtime-2.1. On the Qwen side too, the cheapest model, 3.0 Flash, is a different model from the 3.1 whose performance was compared earlier.
Combining the old Flash's pricing with the new 3.1's performance results would create a price-performance ratio that doesn't actually exist.
APIs specialized for transcription or speech generation are easier to understand if viewed separately from real-time conversation APIs.
| Dedicated API | Region | Input per 1M tokens | Output per 1M tokens |
|---|---|---|---|
| Qwen-Audio-3.1-ASR-Flash | Singapore | $0.15 | $0.47 |
| Qwen-Audio-3.1-TTS-Flash | Singapore | $0.23 | $1.87 |
| Qwen-Audio-3.1-TTS-Next | Beijing | $0.848 | $1.696 |
Based on the ASR model pricing, the TTS-Flash revision notice and the TTS-Next specifications.
ASR generates text from audio and TTS generates audio, so this is not a table for deciding "which API is cheaper" by comparing output prices alone. TTS-Next is also offered in a different region.
The details of the price cuts can be confirmed concretely.
In the price revision effective September 22, the text + audio output of the international Qwen-Audio-3.0-Realtime-Flash was cut from $15 to $1.87 per million tokens.
Calculated as "1 − 1.87 ÷ 15," the reduction is about 87.5%. But this applies to a specific output tier of the old Flash, and does not mean the 3.1 Plus price fell by the same proportion.
TTS-Flash, meanwhile, moved from a pricing scheme of $0.15 per 10,000 input characters to billing based on the number of input and output tokens.
Converting characters into the length of finished audio varies with the content of the script and speaking speed. Because the billing unit itself has changed, comparing the actual invoice for generating the same script is more practical than calculating a uniform reduction rate.
For real-time conversation APIs too, companies count audio tokens and handle conversation history differently, so a price difference per million tokens does not translate directly into a difference in cost per minute of conversation.
When adopting these services, you also need to check constraints that the price list does not reveal.
According to Qwen's Realtime user guide, tool calling and web search cannot be enabled at the same time.
Separately from the model's own context limit of 262,144 tokens, the audio history it can retain is limited to a maximum of 50 turns and 300 seconds cumulatively, and older history beyond the limit is discarded.
The 300 seconds is not a time limit on the whole call, but it affects how far back past audio can be referenced in long consultations or conversations.
For Japanese-language customer support, you would need to measure not only price but also how often names and proper nouns are misheard, how often the AI wrongly reacts to bystanders' conversation, and how many extra conversational round trips are caused by correcting recognition errors.
For audio content production, consider the number of regenerations caused by misreadings or performance fixes, and their cost.
If a model meets the accuracy your actual use requires and also lowers the cost per finished task or deliverable, the added options in this release could become strong candidates for bringing voice AI into real services.
