On September 23, Google announced two text-to-speech models: "Gemini 3.8 Flash TTS" and "Gemini 3.8 Flash-Lite TTS."

The models let users create new voices by describing their characteristics in text, and adjust emotion and delivery for each line of a script. Flash is positioned for content production that prioritizes expressiveness, while Flash-Lite is aimed at high-volume generation and low latency.

ElevenLabs and Qwen also support instructions for emotion and speaking style, but Google reported strong results for overall voice quality and Japanese naturalness. Looking at individual evaluation categories, however, competitors come out ahead in some areas, and the pricing comes with a condition: it applies only through the end of 2026.

The right model depends on what kind of voice you want to use, and in what kind of work or service.

AD

Creating the voice itself, then adjusting the performance line by line

Google's announcement combines a feature for creating voices with a mechanism for finely controlling how those voices speak.

Users can describe a voice's characteristics in text to create a voice for a person or character, then save it and reuse it repeatedly. Emotion and speaking speed can be specified for each line, and laughter, sighs, and backchannel responses can be added.

According to the model specifications, Flash supports 130 languages and Flash-Lite supports 101, both including Japanese.

Availability has begun in the Gemini API and Google AI Studio, and the models will also roll out to Gemini Notebook (Flash) and Google Vids (Flash-Lite). API access through Gemini Enterprise and a remix feature for adjusting existing voices are planned for later.

When producing content, the way you generate multi-speaker conversations differs between existing voices and voices you create yourself.

According to the API guide, when using existing voices, a single request can generate a conversation of up to two speakers. When combining designed or cloned voices, however, you must generate audio for each speaker separately and then stitch it together.

If you are creating dialogue between original characters, you need to factor in this editing step.

This TTS is also a model for converting text into speech. Its role differs from Gemini Live, which understands what the other person says and works out the content of its replies.

When building it into a voice agent, you need to consider pricing and processing time separately for the model that generates responses and the TTS that converts them into speech.

Near the top overall, but ElevenLabs leads in some categories

In the Hume AI comparison that Google included on page 2 of its evaluation document, Gemini 3.8 Flash TTS scored 0.920 for overall quality and Flash-Lite scored 0.914.

These are high values among the models listed, but the overall score and the individual categories are not on the same scale.

According to Hume's explanation, the overall value is a 0–1 index combining expressiveness and stability, while individual quality categories are rated on a 1–5 scale. Because these are not accuracy rates, they cannot be read as percentages.

Model Overall quality Natural intonation variation Multiple speakers Following a single performance instruction
Gemini 3.8 Flash TTS 0.920 4.58 4.14 4.34
Gemini 3.8 Flash-Lite TTS 0.914 4.51 4.10 4.32
Gemini 3.1 Flash TTS 0.783 3.95 3.60 4.31
ElevenLabs v3 0.706 5.00 3.85 3.89
ElevenLabs v3 Conversational 0.769 4.97 — 3.97
Cartesia Sonic 3.6 0.840 3.40 — 3.37
OpenAI gpt-4o-mini-tts 0.740 4.22 — —
Inworld TTS-2 0.576 3.76 — 4.15

The source is Hume AI's evaluation as published by Google, as of September 2026. Higher is better in all columns. "—" means no score was published, not that the feature is unsupported.

Google explained that for its own models, it used the models as actually offered with default generation settings, and evaluated a single generation result.

One point worth noting is ElevenLabs v3's score for "natural intonation variation."

This category evaluates whether changes in delivery are monotonous, or conversely so exaggerated as to be unnatural for human speech. ElevenLabs v3, which scores below Gemini on overall quality, records 5.00 here.

Meanwhile, Gemini's "following a single performance instruction" score improved only slightly, from 4.31 for the previous version to 4.34.

The large gain in overall quality should not be interpreted as meaning that the ability to handle every kind of performance instruction has improved substantially.

A similar pattern appears in voice creation.

The same document gives an English overall score of 71.4 for Gemini and 70.8 for ElevenLabs Voice Design v3. On voice quality, however, Gemini scores 74.6 and ElevenLabs 76.6, reversing the order.

Because these figures show no margin of error, it is impossible to judge whether such small differences are statistically meaningful.

Whether you can create a voice that suits a work, and whether it can perform as the script instructs, need to be checked separately using actual lines.

AD

In Japanese, Flash leads while Flash-Lite and the previous version are about even

Gemini 3.8 Flash TTS also earns a high rating in the Voice Arena Japanese ranking.

The values confirmed as of September 24 are as follows. Because the numbers have changed since the time they appeared in Google's announcement, we use the latest Voice Arena values here.

Model Japanese Elo Displayed 95% confidence interval
Gemini 3.8 Flash TTS 1209 ±21
Gemini 3.8 Flash-Lite TTS 1152 ±21
Gemini 3.1 Flash TTS 1148 ±10
ElevenLabs Eleven v3 1048 ±10
OpenAI gpt-4o-mini-tts 975 ±10

Elo is a relative rating calculated from head-to-head listening comparisons between models. It does not indicate an accuracy rate or a multiple of quality.

For Flash-Lite and the earlier 3.1 Flash, the rank range for both is 2nd to 3rd.

Looking only at the naturalness of Japanese speech, it cannot be said that Flash-Lite clearly surpasses the previous version.

Under the evaluation method, listeners compare models on 100 sentences prepared for each language and choose which voice sounds more human and natural.

Standard catalog voices are used, and voice cloning is not evaluated. Rankings are also calculated independently for each language, so Japanese Elo cannot be directly compared with English Elo.

These results are useful for narrowing down candidates for Japanese TTS.

However, they do not evaluate response speed, generation cost, or how a system reacts when a user interrupts during real-time conversation.

Even if you consult the same ranking, the conditions you ultimately need to verify differ between producing long-form narration and handling phone calls.

Comparing with Qwen: treat TTS-Flash and TTS-Next separately

In the Qwen Audio 3.1 series, the model that handles ordinary text reading is Qwen-Audio-3.1-TTS-Flash.

It lets users specify emotion and speaking speed, and supports streaming generation and voice cloning.

On features alone, TTS that can be instructed on emotion and performance has already become a major area of competition among vendors.

Model Japanese / language support Main features for audio production Points to check when choosing
Gemini 3.8 Flash TTS 130 languages including Japanese Voice creation, line-level performance, backchannels Conversations with custom voices must be generated per speaker and combined
Gemini 3.8 Flash-Lite TTS 101 languages including Japanese High-volume generation and low-latency processing using performance instructions Check the quality difference from Flash in your actual use case
Qwen-Audio-3.1-TTS-Flash Supports Japanese; supported languages vary by selected voice Emotion and speed specification, voice cloning, streaming You must choose a voice that supports Japanese
Qwen-Audio-3.1-TTS-Next Chinese and English in the public API specification Generates speech together with sound effects and ambient sound Non-streaming. Its purpose differs from ordinary Japanese reading
Eleven v3 74 languages including Japanese Emotion and laughter specification, multi-speaker conversation Check performance fit with your actual voice and script

The Qwen voice list explains that if a selected voice does not support the input language, there may be problems with pronunciation and naturalness.

Even if a model as a whole is described as supporting Japanese, not every voice can be used equally well in Japanese.

Furthermore, the TTS-Next API specification supports generating content that combines speech and sound effects, producing up to 4 minutes of audio at once for podcasts and up to 120 seconds otherwise.

Eleven v3 also emphasizes multi-speaker expression, but it cannot be directly compared unless conditions such as the generation target and supported languages are aligned.

Qwen Audio 3.1 is not included in the Hume AI comparison published by Google or in the Voice Arena Japanese ranking we checked. These results alone therefore cannot determine whether Gemini or Qwen has better sound quality.

AD

Pricing: separate the period through end of 2026 from 2027 onward

Under the standard paid pricing for the Gemini API, audio output costs $9 per million tokens for Flash and $6 for Flash-Lite.

However, these prices apply only through December 31, 2026, and from January 1, 2027 they are scheduled to become $18 and $12, respectively.

  • 2026年9月確認の現行価格
  • 2027年1月からの予定価格
Gemini TTSの音声出力単価横棒グラフ。カテゴリ 3 件、系列: 2026年9月確認の現行価格, 2027年1月からの予定価格(単位: ドル/100万トークン)3.8 Flash TTS3.8 Flash TTS3.8 Flash TTS — 2026年9月確認の現行価格: 9ドル/100万トークン93.8 Flash TTS — 2027年1月からの予定価格: 18ドル/100万トークン183.8 Flash-Lite TTS3.8 Flash-Lite TT…3.8 Flash-Lite TTS — 2026年9月確認の現行価格: 6ドル/100万トークン63.8 Flash-Lite TTS — 2027年1月からの予定価格: 12ドル/100万トークン123.1 Flash TTS3.1 Flash TTS3.1 Flash TTS — 2026年9月確認の現行価格: 20ドル/100万トークン20—単位: ドル/100万トークン
データを表で見る
2026年9月確認の現行価格 (ドル/100万トークン)2027年1月からの予定価格 (ドル/100万トークン)
3.8 Flash TTS918
3.8 Flash-Lite TTS612
3.1 Flash TTS20—
Gemini TTSの音声出力単価標準有料枠の音声出力のみ。3.8の現行価格は2026年末まで。旧3.1の将来価格は表示しない。入力・再生成・割引を含まない。出典: Google Gemini Developer API pricing、2026年9月24日確認

Flash-Lite's audio output price is 70% lower than the previous 3.1 Flash TTS through the end of 2026, and 40% lower even at the planned 2027 price.

These figures use the previous version's current price of $20 as the baseline, calculated as 1 − 6 ÷ 20 and 1 − 12 ÷ 20. They are not a comparison that predicts the previous version's price from 2027 onward.

Adding input pricing and Qwen's pricing gives the following.

Model / offering terms Input, per million tokens Output, per million tokens
Gemini 3.8 Flash TTS, through end of 2026 $0.50 $9
Gemini 3.8 Flash-Lite TTS, through end of 2026 $0.50 $6
Gemini 3.8 Flash TTS, planned from 2027 $1.00 $18
Gemini 3.8 Flash-Lite TTS, planned from 2027 $1.00 $12
Qwen-Audio-3.1-TTS-Flash, international site, Singapore $0.23 $1.87

The Qwen figures are based on the price revision notice effective September 22.

All of these are prices for text input and audio output, but the same number of tokens does not necessarily mean the same length of audio.

For that reason, you cannot simply divide the unit prices in the table to conclude that Qwen can generate the same audio some percentage cheaper.

Google states explicitly that it calculates one second of audio as 25 tokens.

At current standard pricing, generating one hour of audio costs $0.81 for Flash and $0.54 for Flash-Lite for audio output alone.

The calculation is 3,600 seconds × 25 ÷ 1 million × output unit price. It does not include input costs or the cost of regeneration. This conversion method also cannot be applied directly to Qwen.

For ongoing use, not just price but also how created voices can be managed becomes important.

Under Google's voice replication API, in addition to a 10–30 second reference recording from the same adult speaker, a separately recorded audio statement from that person showing consent is required.

Up to 200 custom voices can be stored per project, with a validity period of one year. A key used in a mode that does not store the voice on the server expires after seven days.

Google says it embeds a SynthID watermark in generated audio, but this does not remove the need to obtain the person's consent for use of their voice.

Migrating from the older Gemini TTS also takes more than just rewriting the model name.

Performance instructions must be moved from within the script into dedicated metadata, and the default format for non-streaming output has changed from raw PCM to WAV, so existing processing on your side must be updated as well.

If you are making long Japanese narration, can Flash's naturalness reduce corrections and regeneration? If you are doing large-scale reading aloud, can Flash-Lite or Qwen maintain the pronunciation accuracy and voice quality you need?

When deciding what to adopt, it is best to prepare a script you will actually use, generate under the same conditions, and compare the time and cost required to finish.

If voice creation and performance adjustment can be built into a reusable production workflow, it becomes easier to turn the model's expressiveness into real production efficiency.