On September 2, 2026, Google announced "Gemini 3.8 Flash" and a security-focused variant, "Gemini 3.8 Flash Cyber." This marks the third update in the series—three weeks after 3.7 Flash and six weeks after 3.6 Flash. The standard version is generally available at the same limited-time pricing as the previous generation, while the Cyber version is restricted to vetted defensive organizations. Google's own benchmark tables show substantial improvements, but once you factor in Arena's human preference scores, it's still too early to say this model uniformly displaces higher-tier models.
From 65.3% to 73.7% in Three Weeks—Where the Gains Are, and Where the Gaps Remain
In Google's own side-by-side comparison table, 3.8 Flash outperformed 3.7 Flash across several benchmarks: 73.7% vs. 65.3% on DeepSWE v1.1, 61.4% vs. 59.0% on the finance agent benchmark, 10.0% vs. 8.8% on the legal agent benchmark, and 54.9% vs. 53.6% on HLE-Verified. On the other hand, it scored 19.1% on Terminal-bench 4.0—well behind Claude Opus 5's 51.8% and GPT-5.6 Sol's 37.3%—and 59.0% on OSWorld-2.0, trailing Opus 5's 75.4%.
| Benchmark | 3.8 Flash | 3.7 Flash | Table's top score |
|---|---|---|---|
| DeepSWE v1.1 (long-horizon software dev) | 73.7% | 65.3% | Claude Opus 5: 74.0% |
| Vals Finance Agent v2 | 61.4% | 59.0% | 3.8 Flash: 61.4% |
| Harvey's Legal Agent Benchmark | 10.0% | 8.8% | 3.8 Flash: 10.0% |
| Terminal-bench 4.0 | 19.1% | 11.2% | Claude Opus 5: 51.8% |
| GDP.pdf | 35.0% | 34.0% | GPT-5.6 Sol: 40.0% |
| CharXiv Reasoning (no tools) | 86.2% | 84.5% | 3.8 Flash: 86.2% |
| HLE-Verified | 54.9% | 53.6% | 3.8 Flash: 54.9% |
| OSWorld-2.0 | 59.0% | 50.6% | Claude Opus 5: 75.4% |
The biggest gains cluster around multi-step software development tasks carried through to completion, financial and legal document processing, and understanding of charts and long-form video. On the harder BioMysteryBench problem set, scores rose from 43.5% to 56.5%, and LABBench2 climbed from 82.1% to 86.2%. By contrast, on Terminal-bench 4.0—which measures general-purpose agentic ability—the model gained 7.9 points over the previous generation but still trails Opus 5 by 32.7 points. Even when benchmark names sound similar, terminal operation, long-horizon development, and OS control are distinct capabilities; a single figure like 73.7% cannot be generalized into overall superiority in software development work.
The comparison table itself comes with caveats. Google places its own Gemini scores alongside self-reported or publicly disclosed figures from other companies in a single table, and there's no guarantee every model was re-run by the same organization under identical API settings. Even 3.8 Flash's 54.9% on HLE-Verified is only 0.4 points ahead of GPT-5.6 Sol's 54.5% and 0.5 points ahead of Opus 5's 54.4%. Given margins this small, what matters more than the ranking is confirming exactly which task set and reasoning configuration produced each number.
Arena's 1494 Is Provisional—A 3-Point Gap Isn't a Verdict
Arena has users pick their preferred response from two anonymized answers, which yields a signal closer to real-world usage than accuracy on a fixed test set. The "rank spread" that Arena reports alongside each score accounts for models whose confidence intervals overlap, showing the range of ranks a model could plausibly occupy. The raw rank is the current best estimate, but it doesn't settle superiority against rivals whose rank spread overlaps with it.
The figures used here come from the aggregation with Arena's default style-control adjustment applied. In the Text Arena overall leaderboard for September 2, 2026, Gemini 3.8 Flash (High) held a provisional 7th place with a score of 1494±9, a rank spread of 2nd–21st, and 5,137 votes. Gemini 3.7 Flash (High) ranked 10th with 1491±8—a gap of just 3 points, with overlapping confidence intervals. At this stage, Arena's data does not support a clear win for 3.8.
| Text Arena Overall (Sep 2) | Score | Raw Rank | Rank Spread | Votes |
|---|---|---|---|---|
| Claude Fable 5 | 1508±5 | 1st | 1st–4th | 27,013 |
| Claude Opus 4.6 High | 1505±4 | 2nd | 1st–5th | 72,114 |
| Gemini 3.8 Flash High | 1494±9 | 7th | 2nd–21st | 5,137 |
| Gemini 3.7 Flash High | 1491±8 | 10th | 3rd–23rd | 5,682 |
The "provisional" label exists for a reason. Arena tests some models anonymously before public release, and keeps their standing marked as provisional until enough post-release votes accumulate. The rough threshold for full listing is 1,000 votes, but for models that accumulated votes before release, Arena still needs to confirm the score holds up once evaluated under post-release usage conditions. 3.8 Flash already has 5,137 votes, yet its rank spread of 2nd to 21st remains wide. Google's fixed benchmarks show a clear generational gap; Arena shows a small, uncertain one. These aren't contradictory—they're simply measuring different things.
At the Same $0.75, Reasoning Volume Is What Drives Total Cost
3.8 Flash's Standard pricing runs through December 31, 2026, at $0.75 per million input tokens and $3.75 per million output tokens (including thinking tokens)—identical to 3.7 Flash. On January 1, 2027, those rates rise to $1.50 and $7.50 respectively. In a simple scenario using one million tokens each of input and output, the total cost would double from $4.50 to $9.00.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Simple total (1M in + 1M out) |
|---|---|---|---|
| Gemini 3.8 Flash (through year-end) | $0.75 | $3.75 | $4.50 |
| Claude Opus 5 | $5.00 | $25.00 | $30.00 |
| Claude Sonnet 5 | $2.00 | $10.00 | $12.00 |
| GPT-5.6 Sol | $4.00 | $20.00 | $24.00 |
| GPT-5.6 Terra | $2.00 | $12.00 | $14.00 |
For equivalent workloads, 3.8 Flash's current-year pricing undercuts both Opus 5 and Sonnet 5. But Google itself notes that on complex tasks, 3.8 Flash may take additional reasoning steps, call tools repeatedly, and consume more tokens under higher reasoning effort. Even switching from 3.7 at an identical per-token rate won't produce the same bill if output length, retries, and tool calls increase. What teams should compare when adopting the model isn't the price per million tokens, but the total tokens, latency, and retry count needed to complete the same job.
The API caps stand at 1,048,576 input tokens and 65,536 output tokens, and the model is available under the Stable identifier gemini-3.8-flash. It accepts text, image, and video input, as well as audio and PDFs. It supports code execution and function calling. Google Search and file search are available, as is structured output. Computer use is in Preview; Live API, image generation, and audio generation are not supported. The model card lists hallucination and occasional latency or timeouts as known limitations, and its knowledge cutoff varies by domain—March 2026 for some areas, January 2025 for others.
The Cyber Variant's 86.2% Isn't the Performance You Get From the General API
Gemini 3.8 Flash Cyber posted an 86.2% pass@1 rate on CyberGym, beating 3.5 Flash Cyber's 77.5% by 8.7 points and GPT-5.6 Sol's 83.6% by 2.6 points. On CWE-Bench, it scored 47.2%, just 0.6 points short of Fable 5's 47.8%, though Google frames this as Pareto-optimal once cost is factored in.
CyberGym evaluates the ability to discover vulnerabilities in C/C++ code, while CWE-Bench measures the ability to fix them. Google says its internal evaluation spanning 20 languages showed a success rate above 70%, and that on Chrome Security tasks, the model produced correct fixes at 2.6 times the rate of large commercial models. Wiz also reported that in its own internal penetration testing, the model achieved a 7.5–9.7 percentage point higher recall rate at 2.3 to 5.2 times lower cost. However, none of these task sets or comparison baselines have been made public. Without absolute figures or raw data disclosed, these results cannot be read as independently reproducible.
Access conditions are stricter than for the standard version. Google's Fairwind Program lists over 650 participating organizations, but restricts use to internal security, incident response, and penetration testing personnel, and requires phishing-resistant multi-factor authentication and access tracking. Sharing, redistributing, or reselling model access is also prohibited. No public pricing has been disclosed for the Cyber version, and the standard version's $0.75 input / $3.75 output rates cannot simply be applied to it. That 86.2% figure reflects a restricted model run under restricted conditions—it is not the cybersecurity performance you'd get by calling 3.8 Flash through the general Gemini API.
Why the Official Terminal-bench Figures Don't Match Remains Unexplained
On Terminal-bench 2.1, Google DeepMind's general-audience comparison table lists 89.4% for 3.8 Flash and 85.8% for 3.7 Flash, while the Gemini Enterprise Agent Platform's developer guide lists 90.8% and 81.6% for the same models. Both pages report different figures under the identical benchmark name, and neither publicly explains the difference in configuration or timing, making it impossible to pin down the generational improvement to a single number.
The discrepancy is not a minor rounding error. Calculated from the first set of figures, the improvement comes to 3.6 points; from the second, 9.2 points. There's no basis for declaring either set wrong, and it would be inappropriate to combine the higher 3.8 figure with the lower 3.7 figure to maximize the apparent gap. Until Google updates its methodology or reconciles the tables, readers should stick to comparisons drawn from within a single source.
On fixed benchmarks, 3.8 Flash has narrowed the gap between the Flash price tier and higher-end models. But Arena's provisional scores, the way reasoning volume drives total cost, the Cyber version's access restrictions, and the mismatch in official figures don't disappear just because a leaderboard was published on launch day. Teams deciding whether to adopt the model should run 3.7 and 3.8 side by side on identical tasks, identical reasoning settings, and identical harnesses during the current same-price window through year-end. By logging completion rates and total token counts—along with latency and retry counts—teams can determine for themselves, using their own numbers, whether the migration's benefits hold up once prices double in January 2027.
