On September 23, Google announced two text-to-speech models: "Gemini 3.8 Flash TTS" and "Gemini 3.8 Flash-Lite TTS."
The models let users create new voices by describing their characteristics in text, and adjust emotion and delivery for each line of a script. Flash is positioned for content production that prioritizes expressiveness, while Flash-Lite is aimed at high-volume generation and low latency.
ElevenLabs and Qwen also support instructions for emotion and speaking style, but Google reported strong results for overall voice quality and Japanese naturalness. Looking at individual evaluation categories, however, competitors come out ahead in some areas, and the pricing comes with a condition: it applies only through the end of 2026.
The right model depends on what kind of voice you want to use, and in what kind of work or service.
Creating the voice itself, then adjusting the performance line by line
Google's announcement combines a feature for creating voices with a mechanism for finely controlling how those voices speak.
Users can describe a voice's characteristics in text to create a voice for a person or character, then save it and reuse it repeatedly. Emotion and speaking speed can be specified for each line, and laughter, sighs, and backchannel responses can be added.
According to the model specifications, Flash supports 130 languages and Flash-Lite supports 101, both including Japanese.
Availability has begun in the Gemini API and Google AI Studio, and the models will also roll out to Gemini Notebook (Flash) and Google Vids (Flash-Lite). API access through Gemini Enterprise and a remix feature for adjusting existing voices are planned for later.
When producing content, the way you generate multi-speaker conversations differs between existing voices and voices you create yourself.
According to the API guide, when using existing voices, a single request can generate a conversation of up to two speakers. When combining designed or cloned voices, however, you must generate audio for each speaker separately and then stitch it together.
If you are creating dialogue between original characters, you need to factor in this editing step.
This TTS is also a model for converting text into speech. Its role differs from Gemini Live, which understands what the other person says and works out the content of its replies.
When building it into a voice agent, you need to consider pricing and processing time separately for the model that generates responses and the TTS that converts them into speech.
Near the top overall, but ElevenLabs leads in some categories
In the Hume AI comparison that Google included on page 2 of its evaluation document, Gemini 3.8 Flash TTS scored 0.920 for overall quality and Flash-Lite scored 0.914.
These are high values among the models listed, but the overall score and the individual categories are not on the same scale.
According to Hume's explanation, the overall value is a 0–1 index combining expressiveness and stability, while individual quality categories are rated on a 1–5 scale. Because these are not accuracy rates, they cannot be read as percentages.
| Model | Overall quality | Natural intonation variation | Multiple speakers | Following a single performance instruction |
|---|---|---|---|---|
| Gemini 3.8 Flash TTS | 0.920 | 4.58 | 4.14 | 4.34 |
| Gemini 3.8 Flash-Lite TTS | 0.914 | 4.51 | 4.10 | 4.32 |
| Gemini 3.1 Flash TTS | 0.783 | 3.95 | 3.60 | 4.31 |
| ElevenLabs v3 | 0.706 | 5.00 | 3.85 | 3.89 |
| ElevenLabs v3 Conversational | 0.769 | 4.97 | — | 3.97 |
| Cartesia Sonic 3.6 | 0.840 | 3.40 | — | 3.37 |
| OpenAI gpt-4o-mini-tts | 0.740 | 4.22 | — | — |
| Inworld TTS-2 | 0.576 | 3.76 | — | 4.15 |
The source is Hume AI's evaluation as published by Google, as of September 2026. Higher is better in all columns. "—" means no score was published, not that the feature is unsupported.
Google explained that for its own models, it used the models as actually offered with default generation settings, and evaluated a single generation result.
One point worth noting is ElevenLabs v3's score for "natural intonation variation."
This category evaluates whether changes in delivery are monotonous, or conversely so exaggerated as to be unnatural for human speech. ElevenLabs v3, which scores below Gemini on overall quality, records 5.00 here.
Meanwhile, Gemini's "following a single performance instruction" score improved only slightly, from 4.31 for the previous version to 4.34.
The large gain in overall quality should not be interpreted as meaning that the ability to handle every kind of performance instruction has improved substantially.
A similar pattern appears in voice creation.
The same document gives an English overall score of 71.4 for Gemini and 70.8 for ElevenLabs Voice Design v3. On voice quality, however, Gemini scores 74.6 and ElevenLabs 76.6, reversing the order.
Because these figures show no margin of error, it is impossible to judge whether such small differences are statistically meaningful.
Whether you can create a voice that suits a work, and whether it can perform as the script instructs, need to be checked separately using actual lines.
In Japanese, Flash leads while Flash-Lite and the previous version are about even
Gemini 3.8 Flash TTS also earns a high rating in the Voice Arena Japanese ranking.
The values confirmed as of September 24 are as follows. Because the numbers have changed since the time they appeared in Google's announcement, we use the latest Voice Arena values here.
| Model | Japanese Elo | Displayed 95% confidence interval |
|---|---|---|
| Gemini 3.8 Flash TTS | 1209 | ±21 |
| Gemini 3.8 Flash-Lite TTS | 1152 | ±21 |
| Gemini 3.1 Flash TTS | 1148 | ±10 |
| ElevenLabs Eleven v3 | 1048 | ±10 |
| OpenAI gpt-4o-mini-tts | 975 | ±10 |
Elo is a relative rating calculated from head-to-head listening comparisons between models. It does not indicate an accuracy rate or a multiple of quality.
For Flash-Lite and the earlier 3.1 Flash, the rank range for both is 2nd to 3rd.
Looking only at the naturalness of Japanese speech, it cannot be said that Flash-Lite clearly surpasses the previous version.
Under the evaluation method, listeners compare models on 100 sentences prepared for each language and choose which voice sounds more human and natural.
Standard catalog voices are used, and voice cloning is not evaluated. Rankings are also calculated independently for each language, so Japanese Elo cannot be directly compared with English Elo.
These results are useful for narrowing down candidates for Japanese TTS.
However, they do not evaluate response speed, generation cost, or how a system reacts when a user interrupts during real-time conversation.
Even if you consult the same ranking, the conditions you ultimately need to verify differ between producing long-form narration and handling phone calls.
Comparing with Qwen: treat TTS-Flash and TTS-Next separately
In the Qwen Audio 3.1 series, the model that handles ordinary text reading is Qwen-Audio-3.1-TTS-Flash.
It lets users specify emotion and speaking speed, and supports streaming generation and voice cloning.
On features alone, TTS that can be instructed on emotion and performance has already become a major area of competition among vendors.
| Model | Japanese / language support | Main features for audio production | Points to check when choosing |
|---|---|---|---|
| Gemini 3.8 Flash TTS | 130 languages including Japanese | Voice creation, line-level performance, backchannels | Conversations with custom voices must be generated per speaker and combined |
| Gemini 3.8 Flash-Lite TTS | 101 languages including Japanese | High-volume generation and low-latency processing using performance instructions | Check the quality difference from Flash in your actual use case |
| Qwen-Audio-3.1-TTS-Flash | Supports Japanese; supported languages vary by selected voice | Emotion and speed specification, voice cloning, streaming | You must choose a voice that supports Japanese |
| Qwen-Audio-3.1-TTS-Next | Chinese and English in the public API specification | Generates speech together with sound effects and ambient sound | Non-streaming. Its purpose differs from ordinary Japanese reading |
| Eleven v3 | 74 languages including Japanese | Emotion and laughter specification, multi-speaker conversation | Check performance fit with your actual voice and script |
The Qwen voice list explains that if a selected voice does not support the input language, there may be problems with pronunciation and naturalness.
Even if a model as a whole is described as supporting Japanese, not every voice can be used equally well in Japanese.
Furthermore, the TTS-Next API specification supports generating content that combines speech and sound effects, producing up to 4 minutes of audio at once for podcasts and up to 120 seconds otherwise.
Eleven v3 also emphasizes multi-speaker expression, but it cannot be directly compared unless conditions such as the generation target and supported languages are aligned.
Qwen Audio 3.1 is not included in the Hume AI comparison published by Google or in the Voice Arena Japanese ranking we checked. These results alone therefore cannot determine whether Gemini or Qwen has better sound quality.
Pricing: separate the period through end of 2026 from 2027 onward
Under the standard paid pricing for the Gemini API, audio output costs $9 per million tokens for Flash and $6 for Flash-Lite.
However, these prices apply only through December 31, 2026, and from January 1, 2027 they are scheduled to become $18 and $12, respectively.
- 2026年9月確認の現行価格
- 2027年1月からの予定価格
データを表で見る
| 2026年9月確認の現行価格 (ドル/100万トークン) | 2027年1月からの予定価格 (ドル/100万トークン) | |
|---|---|---|
| 3.8 Flash TTS | 9 | 18 |
| 3.8 Flash-Lite TTS | 6 | 12 |
| 3.1 Flash TTS | 20 | — |
Flash-Lite's audio output price is 70% lower than the previous 3.1 Flash TTS through the end of 2026, and 40% lower even at the planned 2027 price.
These figures use the previous version's current price of $20 as the baseline, calculated as 1 − 6 ÷ 20 and 1 − 12 ÷ 20. They are not a comparison that predicts the previous version's price from 2027 onward.
Adding input pricing and Qwen's pricing gives the following.
| Model / offering terms | Input, per million tokens | Output, per million tokens |
|---|---|---|
| Gemini 3.8 Flash TTS, through end of 2026 | $0.50 | $9 |
| Gemini 3.8 Flash-Lite TTS, through end of 2026 | $0.50 | $6 |
| Gemini 3.8 Flash TTS, planned from 2027 | $1.00 | $18 |
| Gemini 3.8 Flash-Lite TTS, planned from 2027 | $1.00 | $12 |
| Qwen-Audio-3.1-TTS-Flash, international site, Singapore | $0.23 | $1.87 |
The Qwen figures are based on the price revision notice effective September 22.
All of these are prices for text input and audio output, but the same number of tokens does not necessarily mean the same length of audio.
For that reason, you cannot simply divide the unit prices in the table to conclude that Qwen can generate the same audio some percentage cheaper.
Google states explicitly that it calculates one second of audio as 25 tokens.
At current standard pricing, generating one hour of audio costs $0.81 for Flash and $0.54 for Flash-Lite for audio output alone.
The calculation is 3,600 seconds × 25 ÷ 1 million × output unit price. It does not include input costs or the cost of regeneration. This conversion method also cannot be applied directly to Qwen.
For ongoing use, not just price but also how created voices can be managed becomes important.
Under Google's voice replication API, in addition to a 10–30 second reference recording from the same adult speaker, a separately recorded audio statement from that person showing consent is required.
Up to 200 custom voices can be stored per project, with a validity period of one year. A key used in a mode that does not store the voice on the server expires after seven days.
Google says it embeds a SynthID watermark in generated audio, but this does not remove the need to obtain the person's consent for use of their voice.
Migrating from the older Gemini TTS also takes more than just rewriting the model name.
Performance instructions must be moved from within the script into dedicated metadata, and the default format for non-streaming output has changed from raw PCM to WAV, so existing processing on your side must be updated as well.
If you are making long Japanese narration, can Flash's naturalness reduce corrections and regeneration? If you are doing large-scale reading aloud, can Flash-Lite or Qwen maintain the pronunciation accuracy and voice quality you need?
When deciding what to adopt, it is best to prepare a script you will actually use, generate under the same conditions, and compare the time and cost required to finish.
If voice creation and performance adjustment can be built into a reusable production workflow, it becomes easier to turn the model's expressiveness into real production efficiency.
