On September 21, 2026, xAI officially released Grok 4.7, its latest reasoning model optimized for coding and advanced knowledge work. The update comes just over a month after the debut of the previous generation, Grok 4.6, and maintains the same pricing as its predecessor: $2 per million input tokens and $6 per million output tokens. Alongside the announcement, the model became immediately available through xAI's own Cursor, Grok Build, and official API. On the same day, GitHub also began a phased rollout of Grok 4.7 across all GitHub Copilot plans—Pro, Pro+, Max, Business, and Enterprise. This established a distribution pipeline that allows the model to instantly permeate real-world development environments through a central developer-tools platform.
Disruptive Pricing: $2/$6 per Million Tokens
Grok 4.7's API pricing is set at $2 per million input tokens and $6 per million output tokens. This represents a substantial discount compared to GPT-5.6 Sol ($4 input / $20 output)—half the input cost and roughly a third of the output cost—and an even steeper discount against Claude Fable 5.1 ($10 input / $50 output), at one-fifth the input cost and roughly one-eighth the output cost. A fast variant offering double the standard output speed is also available at double the price ($4 input / $12 output).
API Pricing Comparison Among Leading Frontier Reasoning Models
| Model | Input Price (per 1M tokens) | Output Price (per 1M tokens) | vs. Grok 4.7 (Input) | vs. Grok 4.7 (Output) |
|---|---|---|---|---|
| Grok 4.7 | $2 | $6 | Baseline | Baseline |
| GPT-5.6 Sol | $4 | $20 | 2x | 3.3x |
| Claude Fable 5.1 | $10 | $50 | 5x | 8.3x |
- Input Price
- Output Price
データを表で見る
| Input Price (USD) | Output Price (USD) | |
|---|---|---|
| Grok 4.7 | 2 | 6 |
| GPT-5.6 Sol | 4 | 20 |
| Claude Fable 5.1 | 10 | 50 |
As this comparison shows, Grok 4.7's output-token pricing undercuts other frontier models by a factor of 3.3x to 8.3x, substantially improving the cost-efficiency of agentic operations that rely on lengthy chains of reasoning.
In software development automation, agentic workflows—where reasoning models think through problems across multiple steps—tend to consume tokens at a dramatically accelerated rate. Because API costs balloon rapidly as output tokens accumulate, the expense of running top-tier models continuously has posed a significant constraint for many organizations. By offering pricing on par with low-cost Chinese models, Grok 4.7 gives development teams a way to run autonomous agents frequently while keeping costs under control.
Advances in Model Architecture and Reasoning Training
According to xAI's announcement materials, Grok 4.7 is built on a new, larger base model compared to Grok 4.6. The training process involved extended reinforcement learning (RL) with heavy emphasis on difficult problems that take hours to complete.
This additional training improved the model's ability to verify intermediate results during generation and self-correct errors. Its ability to maintain consistent output over long contexts has also been enhanced. Furthermore, xAI pre-trained Grok 4.7 to natively understand the internal structure of its proprietary autonomous coding environment, the "Grok Bot" harness. This enables more accurate command generation and situational judgment in practical work involving external tool calls and interactive file editing.
Instant Penetration into Development Platforms: Simultaneous Rollout Across All GitHub Copilot Plans
A standout feature of Grok 4.7's distribution strategy is its immediate integration into major development platforms beyond xAI's own ecosystem. In step with its release on Cursor, Grok Build, and the official API, GitHub announced via its official Changelog that Grok 4.7 would be available on GitHub Copilot.
The rollout covers all plans, including Pro and Pro+ for individual users, Max for power users, and Business and Enterprise for organizations. Developers can now select Grok 4.7 directly from the model picker in Visual Studio Code, Visual Studio, JetBrains IDEs, Xcode, Eclipse, as well as the GitHub Copilot CLI, cloud agents, and mobile apps.
Enterprise management features were also in place from day one. Administrators of Copilot Business and Enterprise plans can control access to Grok 4.7 through the model policy screen in organization settings. By default, new models are automatically enabled, with usage billed on a pay-as-you-go basis according to the provider's list price. The fact that the model was built into developers' everyday workflows as an available option from day one is likely to significantly accelerate its adoption.
Official Benchmarks vs. Third-Party Evaluations: A Gap in Agentic Reasoning
A close look at the benchmark results reveals Grok 4.7's strengths and remaining challenges in concrete numbers. On the official CursorBench 4.0 benchmark, Grok 4.7 scored 46.3%, outperforming GPT-5.6 Sol (41.7%). However, on Terminal-Bench 4.0, which evaluates autonomous terminal operation, it scored only 38.0%—a 19.9-point gap behind Fable 5.1 (57.9%).
On DeepSWE v1.1 (high reasoning-effort setting), which simulates real-world software engineering tasks, Grok 4.7 scored 71.0%, surpassing both Grok 4.6 (65.2%) and Claude Fable 5.1 Max (70.0%), though trailing GPT-5.6 Sol Max (72.7%). On EEBench, which covers electrical engineering tasks, it reached 64.0%, well ahead of GPT-5.6 Sol Max (39.4%) and Fable 5.1 Max (56.4%). In specialized domains such as AA Briefcase v1.1, which simulates extended office work, and the Harvey Legal Agent Benchmark for legal tasks, Grok 4.7 also showed substantial improvement over its predecessor.
Key Benchmark Score Comparison
| Benchmark | Evaluation Target | Grok 4.7 | GPT-5.6 Sol Max | Claude Fable 5.1 Max |
|---|---|---|---|---|
| CursorBench 4.0 | Extended coding | 46.3% | 41.7% | 51.8% |
| DeepSWE v1.1 | Software engineering | 71.0% | 72.7% | 70.0% |
| EEBench | Electrical engineering | 64.0% | 39.4% | 56.4% |
| Terminal-Bench 4.0 | Autonomous terminal operation (official) | 38.0% | - | 57.9% |
- Grok 4.7
- GPT-5.6 Sol Max
- Claude Fable 5.1 Max
データを表で見る
| Grok 4.7 (%) | GPT-5.6 Sol Max (%) | Claude Fable 5.1 Max (%) | |
|---|---|---|---|
| CursorBench 4.0 | 46.3 | 41.7 | 51.8 |
| DeepSWE v1.1 | 71 | 72.7 | 70 |
| EEBench | 64 | 39.4 | 56.4 |
| Terminal-Bench 4.0 | 38 | — | 57.9 |
As the chart and table show, Grok 4.7 delivers strong results in certain domains such as CursorBench 4.0 and EEBench, while leaving a clear gap versus Fable 5.1 on Terminal-Bench 4.0, which measures autonomous terminal operation.
Meanwhile, independent evaluators offer a more cautious assessment. According to Artificial Analysis's composite "Intelligence Index v4.3.2" (an aggregate of 10 evaluations), Grok 4.7 earned an overall score of 46. This places it 7 points behind the joint leaders, Claude Fable 5.1 and GPT-6, both at 53, positioning it in the middle tier of the frontier group. In the organization's independent Terminal-Bench 4.0 test conducted under standardized conditions, Grok 4.7 scored just 26%, trailing far behind GPT-6 Astra (60%) and Claude Fable 5.1 (55%).
These figures suggest that while Grok 4.7 delivers exceptional cost-efficiency for certain coding assistance tasks and routine knowledge work, top-tier models still hold an edge in complex autonomous agent operations involving multi-step shell commands and continuous handling of unexpected errors.
Renewed Safeguard Stack and Future Outlook
To support broader practical deployment, xAI has also revised its approach to safety and cyber defense. Grok 4.7 introduces a completely rebuilt safeguard stack.
The new system is designed to reduce "false refusals"—cases where the AI excessively declines to respond to legitimate requests such as security research or vulnerability assessment—while still guarding against malicious jailbreak attempts and attack assistance. On HackerBench v0.3, a benchmark for evaluating malicious cyber tasks, Grok 4.7 kept the pass-through rate for dangerous dual-use prompts to just 3.3%, while minimizing the blocking of legitimate security work. It also posted the top score of 62.4% on LatchBio's biosafety benchmark, a safety metric for the biological domain. xAI has begun offering select cybersecurity partners invite-only access to red-team capabilities intended for defensive research.
Grok 4.7 represents xAI's clear answer to the cost inflation that has accompanied the push for stronger reasoning capabilities among frontier models. While a capability gap remains compared to the very top autonomous reasoning models, its low pricing—$2 for input and $6 for output per million tokens—combined with its immediate rollout to GitHub Copilot, is steadily reshaping the competitive landscape for practical coding-assistance tools.
