On July 21, 2026, Google announced three models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Both 3.6 Flash and 3.5 Flash-Lite became available the same day, while Cyber will soon be offered on a limited trial basis to governments and trusted partners. In contrast, 3.5 Pro—which Google said at its May Google I/O event would roll out "the following month"—remains in partner testing only. This update puts the cost efficiency and use-case-specific design of the publicly available Flash lineup front and center, while the release timeline for the flagship Pro model remains unclear.
A 17% Token Reduction and $7.50 Output Pricing
The changes in Gemini 3.6 Flash show up not so much in single-response performance as in the computational cost of completing a task. According to Google, the model produces 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, with reductions reaching as much as 65% on DeepSWE. For AI agents that reason through multiple steps and repeatedly call tools, it's not just the length of the final answer but also the number of intermediate attempts that drives up both billing and latency. Google says 3.6 Flash also reduced the number of reasoning steps and tool calls required.
Standard API pricing is $1.50 per million input tokens and $7.50 per million output tokens. Compared to 3.5 Flash's $1.50 input and $9.00 output pricing, this keeps input pricing unchanged while cutting output pricing by 16.7%. If you simply combine the 17% token reduction reported on the Artificial Analysis Index with this pricing difference, the output-billed portion of the same evaluation task would work out to roughly a 31% reduction. However, this is a rough calculation based on token volumes from a specific evaluation, not a total cost figure that includes input, caching, search, or external tool expenses.
Even as it lowered prices, 3.6 Flash outperformed 3.5 Flash on Google's evaluations. DeepSWE, which measures software development capability, rose from 37% to 49%, while MLE Bench, which measures machine learning research capability, climbed from 49.7% to 63.9%. OSWorld-Verified, which measures computer operation, also improved from 78.4% to 83.0%, and GDPval-AA v2, which evaluates knowledge work, rose from 1349 to 1421. In the Gemini API and Gemini Enterprise, Computer Use—which handles screen operations—is also available as a client-side built-in tool.

That said, evaluation scores depend on the execution environment of each benchmark. The 3.6 Flash model card indicates a maximum input of 1 million tokens, a maximum output of 64,000 tokens, and a knowledge cutoff of March 2026, while also listing hallucinations along with occasional delays and timeouts as known limitations. 3.6 Flash has begun rolling out across the Gemini app, Gemini API, Google AI Studio, Gemini Enterprise, Google Antigravity, and other platforms.
Lite Isn't Cheaper Than Its Predecessor
Gemini 3.5 Flash-Lite targets low latency and high throughput, aiming for 350 output tokens per second. Pricing is $0.30 per million input tokens and $2.50 per million output tokens—one-fifth the input cost and one-third the output cost of 3.6 Flash. For high-call-volume work like document processing, search, and classification, this low absolute pricing matters.
However, this isn't a price cut from the immediately preceding 3.1 Flash-Lite. Standard pricing for 3.1 Flash-Lite was $0.25 for input and $1.50 for output, meaning 3.5 Flash-Lite is 20% more expensive on input and 66.7% more expensive on output. Google's characterization of it as the "most cost-effective model in the 3.5 class" is a statement that includes comparisons against the mainline Flash model and performance gains—a different character from 3.6 Flash, where the generational update itself lowered billing rates.
While pricing went up, agentic performance also improved substantially. Compared to 3.1 Flash-Lite, Terminal-Bench 2.1 rose from 31% to 54%, the long-context GDM-MRCR v2 climbed from 60.1% to 72.2%, and GDPval-AA v2 went from 642 to 1140. It also outperforms the previous-generation Gemini 3 Flash, one tier up, on SWE-Bench Pro (49.6% vs. 54.2%) and OSWorld-Verified (65.1% vs. 74.0%). Developers can choose to lower the thinking level to prioritize speed and cost, or raise it to hand off multi-step processing.
3.5 Flash-Lite also supports up to 1 million input tokens and up to 64,000 output tokens. It's available on Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform, and will also roll out to the Gemini app and Google Search. For high-volume deployment decisions, organizations need to measure both the per-token price and how many retries are needed to complete a task, using their own data.
Flash Cyber: Calling a Lightweight Model Up to Five Times
Gemini 3.5 Flash Cyber is a model further trained on top of 3.5 Flash specifically for vulnerability discovery, verification, and remediation. Google's CodeMender has multiple Flash Cyber agents examine separate code paths, then consolidates the results into a single report. Rather than calling a single large model once, the design runs cheap models in parallel to widen the scope of exploration.

This distinction is also a condition to keep in mind when reading the published figures. On CyberGym, which deals with real-world software vulnerabilities, CodeMender called Flash Cyber up to five times per final report. The comparison figures from other companies are self-reported and aren't based on running the models alone under a common, shared calling condition. Google's characterization of the results as "frontier-level" applies to the combined system of Flash Cyber and CodeMender together, not the model alone.
On the other hand, in an evaluation that examined the V8 JavaScript engine with a fixed number of calls, Flash Cyber found 55 confirmed unique issues. Standard 3.5 Flash found 47, and Claude Opus 4.6 found 36, with 10 issues found only by Flash Cyber. Whether running a lightweight model repeatedly pays off depends on whether it can broaden exploration across different code paths, rather than repeatedly reporting the same issue.
Because it could also be repurposed for attacks, Flash Cyber will not be made publicly available. Google plans to offer it soon, exclusively to governments and trusted partners, as a pilot delivered via CodeMender. The core CodeMender functionality, which is being made broadly available on the Gemini Enterprise Agent Platform, runs on standard Gemini models—so having access to the same product name doesn't necessarily mean access to Flash Cyber itself.
3.5 Pro Past Its June Target, as Gemini 4 Pretraining Begins
When Google announced Gemini 3.5 on May 19, it stated explicitly that 3.5 Pro was in internal use and would roll out "the following month." But by the time of this July 21 announcement—past that following month of June—the explanation had shifted to the model being in partner testing, with broad availability coming "once it's ready." No new date was given. Based on official information alone, it's clear the originally stated timeline has already passed.
On July 16, Bloomberg reported, citing people familiar with the matter, that 3.5 Pro was running months behind schedule, with Google taking extra time particularly to improve its coding capabilities. This isn't Google's official reason for the delay, but it does align with the company's own emphasis in May on 3.5 Flash as its "strongest agentic coding model," and with today's continued focus on DeepSWE and Computer Use. While the Pro release is delayed, the latest generation available to developers remains limited to the Flash lineup.
Google also disclosed that it has begun what it calls its most ambitious pretraining effort to date, aimed at Gemini 4. No progress figures or release timeline were shared, and this doesn't mean 3.5 Pro will be skipped. Before Gemini 4 arrives, the key indicators to watch are 3.5 Pro's general availability date, its model card, and its API pricing. Only once those three things are known will it be possible to judge whether Google has managed to bring the cost efficiency it achieved with Flash to its flagship model as well.
