On September 30, Google announced its next-generation flagship AI model, Gemini 4 Argon. It raises the maximum number of tokens the model can generate in a single run from roughly 64,000 to 1 million. Google is positioning it as a model that can handle long-running software development and complex professional work such as finance and legal tasks.
In Google's published comparisons, Argon outperformed models from OpenAI and Anthropic on evaluations covering financial research, legal work, and business automation. In software development and scientific research, however, competing models come out ahead on some evaluations.
Google has also priced the model competitively, but access will begin with trusted cyber defenders. It has not said when the model will be available to general developers and users. Taken together, the performance, pricing, and rollout approach show Google trying to strengthen its competitiveness in enterprise professional work and long-running AI agent tasks.
A 1 million-token output limit for long-running work
Argon raises its output limit to 1 million tokens.
A token is the unit an AI uses to process text and code. What has grown substantially is not the amount of information the model can read as input, but the maximum number of tokens it can generate in a single run. Compared with the previous limit of roughly 64,000 tokens, that is about 15.6 times larger.
Google says this headroom lets the model sustain reasoning and generation on the scale of hundreds of thousands of tokens within a single run, so it can work on complex problems for longer.
A large-scale code migration, for example, requires repeating a cycle many times: examining existing code, modifying it, checking test results, and rewriting. A higher output limit is meant to make it easier to carry out such long tasks without being cut off midway.
That said, it does not mean the model will always produce a 1 million-token response, and reasoning for longer does not guarantee reaching the right answer. How long a single run can continue and whether the model actually completes the work correctly need to be considered separately.
According to Google, thousands of its employees are already using Argon for specialized coding, research, and other work.
External availability will be phased in, however. Access will first be offered to trusted cyber defenders through the "Fairwind Program," and will then expand, starting with paid API users and Google AI Ultra subscribers.
As of the announcement, general developers cannot yet use it immediately through the API.
Strong scores in finance and legal; results split by evaluation in development and science
In the comparisons Google published, Argon posts particularly strong results in areas such as financial research, legal work, and business automation.
This suggests the model was developed with an emphasis on work that goes beyond answering one-off questions, involving multiple steps to produce a final deliverable.
| Evaluation and main focus | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1: long-running development work on real codebases | 77.9% | 74.1% | 74.2% |
| FrontierSWE v2: software development | 55.0% | 65.5% | 62.3% |
| Terminal-Bench 4.0: multi-step work in the terminal | 57.4% | 58.2% | 66.4% |
| Terminal-Bench Science 0.1: scientific research tasks | 57.6% | 68.1% | 63.3% |
| LABBench 2 | 88.8% | 85.4% | 73.1% |
| RiemannBench | 76% | 72% | 69.6% |
| AutomationBench: end-to-end execution of business processes | 51.3% | 41.4% | 42.5% |
| Vals Finance Agent v2: multi-step financial research | 65.4% | 53.5% | 58.6% |
| Harvey's Legal Agent Benchmark: legal research and document drafting | 19.6% | 5.4% | 3.8% |
| LVBench: long-form video understanding | 91.7% | 87.5% | 83.7% |
| CWE-bench v1: vulnerability fixing | 68.0% | 68.0% | 67.0% |
The source is Google's evaluation methodology and results. Bold marks the highest value among these three models. Each is an independent metric, so the numbers cannot simply be averaged to judge the overall merit of a model.
On DeepSWE, Argon beat Opus 5.5 by 3.7 points.
On FrontierSWE, another software development evaluation, it trailed Astra by 10.5 points, and on Terminal-Bench 4.0 it trailed Opus 5.5 by 9.0 points.
On Terminal-Bench Science, which covers scientific research, Astra also led by 10.5 points. Yet on LABBench 2 and RiemannBench in the same comparison table, Argon scores highest of the three models.
Even within "software development" or "scientific research," strengths and weaknesses differ depending on the work being evaluated.
The 19.6% in the legal category also calls for caution.
The gap with competing models is large, but it does not mean the model can "automate about 20% of real legal work." It is simply the score on the legal deliverables that benchmark required. The absolute value also shows there is still considerable room before the model can fully handle difficult legal work.
Likewise, the 51.3% on AutomationBench does not mean half of your company's operations can be handed to Argon as is.
The comparison conditions also need attention.
According to Google, Argon was in principle evaluated at its highest thinking setting, measuring whether it can solve a task in a single attempt. Some figures for competing models are taken from the companies' announcements or public leaderboards, so this is not a comparison in which Google re-evaluated every model in the same environment.
On DeepSWE too, Argon used the mini-swe execution environment, while published results were used for the competing models.
Conditions also differ on LVBench, which measures video understanding. Gemini takes one frame per second as input, while Astra uses up to 800 frames and Opus 5.5 uses 600.
On Terminal-Bench Science 0.1, the timeout for the result-checking step was extended to six times the standard setting for Argon.
Given these differences, the numbers are a clue to where each model is strong, but should not be read as a simple overall ranking.
Top of the Vals Index, but running costs vary widely by model
On the "Vals Index v2.1," published by third-party AI evaluator Vals AI, Argon also recorded 68.90%, the highest as of October 1.
Including GPT-6.1 Sol and Claude Sonnet 5.5, which are not in Google's comparison table, reveals differences in both performance and running cost between models.
| Model | Vals Index v2.1 | Displayed cost per test |
|---|---|---|
| Gemini 4 Argon | 68.90% | $15.68 |
| Claude Sonnet 5.5 | 67.04% | $21.34 |
| Claude Opus 5.5 | 66.97% | $32.14 |
| GPT-6 Astra | 63.13% | $18.46 |
| GPT-6.1 Sol | 61.15% | $3.24 |
The source is the Vals AI ranking and methodology, as viewed on October 1. Argon's cost calculation uses the post-introductory-period rates of $4 for input and $20 for output, not the introductory price. Claude's scores also include some tasks that were run on a substitute model after the original model refused to answer because of safety measures.
{
"type": "bar",
"title": "Cost per test on Vals Index v2.1",
"unit": "USD",
"categories": ["Gemini 4 Argon", "Claude Sonnet 5.5", "Claude Opus 5.5", "GPT-6 Astra", "GPT-6.1 Sol"],
"series": [{"name": "Cost displayed by Vals AI", "values": [15.68, 21.34, 32.14, 18.46, 3.24]}],
"caption": "Viewed October 1, 2026. Argon calculated at its regular rates. Depends on evaluation conditions, including thinking amount and use of substitute models; not a direct prediction of production running costs.",
"source": "Vals AI, Vals Index v2.1"
}In this evaluation, Argon scores higher than Opus 5.5 while also costing less per test.
Against GPT-6.1 Sol, however, Argon scores higher but its displayed cost is about 4.8 times as much, which is $15.68 divided by $3.24.
This gap is not determined by API unit prices alone. It also reflects the number of tokens each model used during the evaluation, the length of its reasoning, and execution conditions.
The Vals Index weights results from four fields (finance, coding, law, and tax) by each field's share of US GDP.
It is an attempt to measure how useful AI is in real knowledge work, but 68.90% does not mean the model "can handle 68.9% of the US economy." Nor, of course, does it indicate success rates for work at Japanese companies.
The meaning of the numbers changes greatly depending on which evaluation field your own work most resembles.
Vals also says that if tasks Claude refused are treated as outright failures rather than being run on a substitute model, Sonnet 5.5 scores 65.85% and Opus 5.5 scores 65.05%.
It is worth noting that public rankings reflect not only a model's raw capability but also how it behaves when operated as an actual service.
Introductory pricing undercuts Astra and Opus, and matches Sol and Sonnet 5.5
Argon's introductory price is $2 per million input tokens and $10 per million output tokens.
That is one-fifth of Astra's standard price of $10 input and $50 output, and half of Opus 5.5's $4 input and $20 output.
The price is time-limited, though. Google says it will rise to $4 for input and $20 for output after the introductory period, but has not said when that period ends.
| Model / price tier | Input per 1M tokens | Output per 1M tokens | Cache read per 1M tokens |
|---|---|---|---|
| Argon: introductory price | $2 | $10 | $0.10 |
| Argon: after introductory period | $4 | $20 | Terms after the period to be confirmed |
| GPT-6 Astra | $10 | $50 | $1.00 |
| GPT-6.1 Sol | $2 | $10 | $0.10 |
| Claude Opus 5.5 | $4 | $20 | $0.20 |
| Claude Sonnet 5.5 | $2 | $10 | $0.20 |
The comparison is based on Google's announcement, OpenAI's official pricing, the official Opus 5.5 page, and the Sonnet 5.5 announcement, as of October 1.
It covers base rates for standard processing, and for OpenAI uses Standard pricing at short context lengths. It excludes surcharges for long contexts, fast processing, cache writes and storage, and tool use.
A simple calculation assuming 100,000 input tokens and 20,000 output tokens comes to $0.40 at Argon's introductory price and $0.80 after the introductory period.
Astra comes to $2.00, while Sol and Sonnet 5.5 come to $0.40, the same as Argon's introductory price.
The formula is "input rate × 0.10 + output rate × 0.02," assuming no caching. For Argon's introductory price, for example, that is 2 × 0.10 + 10 × 0.02 = $0.40.
This is only an estimate to illustrate the unit price differences when the same number of tokens is used. It does not mean each model can complete the same job with the same amount of reasoning and output.
On caching, Google says it discounts the regular input token price by 95%. Based on the introductory price, that is $0.10 per million tokens, the same as GPT-6.1 Sol's cache read price.
However, companies count tokens and bill for caching differently. It is hard to judge the total cost of a long-running AI agent from API unit prices alone.
Google's price competition is also not limited to high-priced flagship models.
Argon's introductory price of $2 input and $10 output is on par with GPT-6.1 Sol and Claude Sonnet 5.5. In real deployments, then, a meaningful comparison has to include how many attempts it takes to finish the same job, how long the model reasons, and how much human effort is needed for corrections.
Inside Google: large-scale Rust migration and memory optimization
One of the use cases Google highlighted for Argon is an improved Rust version of the video decoder "libgav1."
An agent using Argon replaced roughly 32,000 lines of SIMD code in the existing Rust port.
SIMD is a technique that speeds up processing by handling multiple pieces of data with a single instruction.
Argon repeatedly measured execution performance and inspected the code generated by the compiler, searching for ways to let the compiler itself vectorize safe Rust code efficiently.
According to Google, without changing the video output, the result was 2.7 times faster than the previous Rust version and approached the performance of the optimized C++ version.
The important point is that the 2.7x figure is measured against the pre-improvement Rust version, not the optimized C++ version.
Google is also using Argon to migrate C/C++ code to Rust, from libraries such as re2 up to the Zircon kernel of Fuchsia OS, which exceeds 800,000 lines.
These large rewrites, however, are still at the stage of automated testing, human audits, and emulation before being put into production. That AI could generate the code is a separate matter from that code already being used in production.
Argon-powered agents also performed analysis and improvements for data center memory optimization.
Google says optimizations that were actually deployed freed more than 300 TiB of memory. It estimates the eventual reduction at 500 TiB to 1 PiB.
The actual reduction of over 300 TiB should be distinguished from the forward-looking estimate of 500 TiB to 1 PiB. Both are also internal results reported by Google, not results reproduced by third parties.
Why cyber defense comes first
Behind the decision to offer Argon first to cyber defense specialists, ahead of general release, is a substantial increase in the model's cybersecurity capabilities.
According to Google, Argon can autonomously discover software vulnerabilities, confirm the problems, and even fix them.
Wiz is also reportedly using Argon in "Scan for Good," which investigates vulnerabilities in critical public infrastructure free of charge.
The ability to find and demonstrate vulnerabilities, however, can be turned to offense as well as defense.
Google is therefore providing Argon to trusted defenders and internal teams with the cybersecurity guardrails removed, while continuing to strengthen safety measures ahead of general availability.
Beyond countering misuse, it says it has also strengthened resistance to "indirect prompt injection," in which instructions embedded in external web pages or documents hijack the AI's behavior.
It will also introduce a mechanism that monitors for behavior in which the model pursues goals in ways that deviate from the user's intent, and stops processing when necessary.
For environments used in high-risk training and evaluation, it is also reinforcing sandboxes isolated from the outside.
Google also participates in a US government initiative under which AI models are voluntarily shared before release. It plans to adjust safety measures and gradually widen access while gathering real-world usage results from early users.
Once access begins for paid API users and Google AI Ultra subscribers, general developers and companies will be able to compare Argon with competing models on their own work.
At that stage, what needs to be measured goes beyond benchmark scores and the 1 million-token limit: the amount of reasoning, processing time, API cost, and human correction effort required to complete real work.
Only then will it be possible to make a practical comparison of which AI to trust with long-running software development and professional work.
