On September 22, 2026, OpenAI expanded its GPT-6 family of frontier reasoning models, making "GPT-6 Sol" (for core development work) and "GPT-6 Luna" (for high-efficiency, high-volume workloads) generally available. Both models inherit the reasoning architecture and alignment techniques of the flagship "GPT-6 Astra," which launched earlier on September 3, 2026. API prices are cut a uniform 50% from the previous generation's promotional pricing.

The models are available immediately through the API as "gpt-6-sol" and "gpt-6-luna." They are also rolling out the same day to Codex, the developer platform, and ChatGPT Work (paid plans including Plus, Pro, and Enterprise). Both models offer a context window of 1,050,000 tokens (about 1.05 million) and keep a maximum output of 128,000 tokens.

The expansion marks an economic turning point in bringing top-tier intelligence to production systems. As autonomous agents spread, the center of gravity in workloads is shifting from single-shot query responses to long-running iterative reasoning and chained execution of external tools. OpenAI's aim is to sharply reduce the total cost of ownership (TCO) per completed task, which is increasingly the deciding factor in model selection.

AD

Two GPT-6 Models for Development Work and High-Volume Processing

The GPT-6 family is a three-tier portfolio with distinct roles. GPT-6 Astra handles highly advanced multi-step reasoning and strategic planning, while the newly released Sol and Luna are designed as the workhorses behind everyday development and high-volume transactions.

GPT-6 Sol is the main model for engineering work that requires a solid level of reasoning and deeper judgment, including exploring complex codebases, multi-step debugging, and precise data extraction. It follows the concise, no-fluff conversational style introduced with Astra and is tuned to avoid excessive jargon and low-value, verbose answers.

GPT-6 Luna, by contrast, pushes speed and cost efficiency to the limit. It targets stable handling of heavy traffic: routine text classification, structured data extraction, summarization, and basic code completion. When the previous-generation GPT-5.6 Luna received an 80% price cut in July 2026, usage grew tenfold, and that track record leads to expectations of very wide adoption for Luna as a routing or preprocessing layer.

On availability, the API, Codex environments, and enterprise workspaces are prioritized. The standard ChatGPT conversation screens on the consumer web and mobile apps are not getting the new models. Luna, however, is being offered first to users of the free desktop app and the Go plan.

How the 50% API Cut Works: Price Table and Effective Rates

The announcement's most direct impact on the developer community is the sweeping revision of standard API pricing. Both GPT-6 Sol and GPT-6 Luna are priced 50% below the promotional prices of the previous-generation GPT-5.6.

  • 入力価格
  • 出力価格
API入力・出力トークン単価比較(100万トークンあたり)横棒グラフ。カテゴリ 4 件、系列: 入力価格, 出力価格(単位: USD)GPT-5.6 SolGPT-5.6 SolGPT-5.6 Sol — 入力価格: 4USD4GPT-5.6 Sol — 出力価格: 20USD20GPT-6 SolGPT-6 SolGPT-6 Sol — 入力価格: 2USD2GPT-6 Sol — 出力価格: 10USD10GPT-5.6 LunaGPT-5.6 LunaGPT-5.6 Luna — 入力価格: 0.2USD0.2GPT-5.6 Luna — 出力価格: 1.2USD1.2GPT-6 LunaGPT-6 LunaGPT-6 Luna — 入力価格: 0.1USD0.1GPT-6 Luna — 出力価格: 0.5USD0.5単位: USD
データを表で見る
入力価格 (USD)出力価格 (USD)
GPT-5.6 Sol420
GPT-6 Sol210
GPT-5.6 Luna0.21.2
GPT-6 Luna0.10.5
API入力・出力トークン単価比較(100万トークンあたり)OpenAI公式発表の料金データに基づく比較出典: OpenAI

As shown above, GPT-6 Sol is priced at $2.00 per million input tokens and $10.00 per million output tokens, exactly half the previous GPT-5.6 Sol ($4.00 input / $20.00 output). GPT-6 Luna costs $0.10 for input and $0.50 for output, a 50% cut on input and a substantial reduction on output compared with GPT-5.6 Luna ($0.20 input / $1.20 output).

Model Input (per 1M tokens) Output (per 1M tokens) Cached input (per 1M tokens) Context length
GPT-6 Sol $2.00 $10.00 $0.20 1,050,000
GPT-5.6 Sol (previous) $4.00 $20.00 - 1,050,000
GPT-6 Luna $0.10 $0.50 $0.01 1,050,000
GPT-5.6 Luna (previous) $0.20 $1.20 - 1,050,000

The repricing significantly changes how enterprises estimate costs for internal systems. Luna's $0.10 input price in particular makes it realistic to run full-text search or preprocessing filters over document archives with millions of items continuously through the API. Sol's pricing of $2.00 input and $10.00 output, meanwhile, encourages always-on deployment of multi-step reasoning agents that teams previously hesitated to adopt on cost grounds.

AD

Cost per Task: AutomationBench and DeepSWE Results

The value of a model is measured less by cheap per-token pricing than by the total cost required to complete a real task successfully. OpenAI published evaluation data showing cost-effectiveness per completed task in tool integration, software development, and computer operation.

On "AutomationBench 1.0.6" (provided by Zapier, a test set that uses 47 integrated tools to measure autonomous execution of business workflows), GPT-6 Sol recorded a 33.2% success rate at its highest reasoning setting (xhigh). The cost was $0.27 per task.

AutomationBench 1.0.6における1タスクあたりの実行コスト横棒グラフ。カテゴリ 4 件、系列: タスク単価(単位: USD/task)GPT-6 Sol (xhigh)GPT-6 Sol (xhigh)GPT-6 Sol (xhigh) — タスク単価: 0.27USD/task0.27GPT-6 Astra (low)GPT-6 Astra (low)GPT-6 Astra (low) — タスク単価: 1.05USD/task1.05Claude Opus 5 (max)Claude Opus 5 (ma…Claude Opus 5 (max) — タスク単価: 3USD/task3Claude Fable 5.1 (max)Claude Fable 5.1 …Claude Fable 5.1 (max) — タスク単価: 2.4USD/task2.4単位: USD/task
データを表で見る
タスク単価 (USD/task)
GPT-6 Sol (xhigh)0.27
GPT-6 Astra (low)1.05
Claude Opus 5 (max)3
Claude Fable 5.1 (max)2.4
AutomationBench 1.0.6における1タスクあたりの実行コスト各モデルの最高または標準実行時におけるAPIコスト実測値出典: OpenAI公表データ

The economic impact of this result is considerable. Compared with the 30.3% ($1.05 per task) achieved by flagship GPT-6 Astra at its low reasoning setting, Sol reaches equal or better accuracy at a much lower cost of $0.27. Against competitor Anthropic's Claude Opus 5 (max setting, 26.9% success rate, $3.00 per task), Sol achieved a higher success rate at one-eleventh the cost (11.1x cheaper).

The same pattern appears in "DeepSWE v1.1," a software development benchmark built on real codebases. GPT-6 Sol (max setting) achieved a 68.8% resolution rate. That is on par with Claude Fable 5 (xhigh setting, 69.9%), while cutting compute cost per task by roughly 80%.

The lightweight GPT-6 Luna also performs notably well. Luna (max setting) achieved a 66.6% resolution rate on DeepSWE v1.1. That means it can complete more than two-thirds of advanced code-fixing tasks at a cost 93% lower than Opus 5 and 96% lower than Fable 5. On "OSWorld 2.0 offline," which tests autonomous operation of computer screens, GPT-6 Sol (xhigh) reached 60.5%, matching the accuracy of Claude Opus 5 (medium, 60.3%) at roughly 80% lower cost per task.

90% Cache Discount and Prefix Protection Improve Operating Efficiency

Alongside lower per-token prices, another factor that shapes real-world TCO is the overhaul of the prompt caching mechanism. For the GPT-6 family, reads of cached input tokens receive a steep 90% discount.

Cached input is processed at just $0.20 per million tokens for Sol and an extremely low $0.01 for Luna. In agent environments where tens of thousands of lines of source code, long API specifications, or detailed system prompts stay fixed at the start of the prompt while the conversation continues, the 90% discount directly reduces running costs.

As a further technical improvement, the cache for the leading prompt prefix is no longer discarded when a developer changes the reasoning effort (thinking level) parameter or when an agent dynamically switches its list of available external tools. By specifying explicit breakpoints, developers can reliably keep the system prompt and base context in cache memory.

The effect of this optimization has already been demonstrated in large-scale operational data from GitHub Copilot, which is under the OpenAI umbrella. An analysis of billions of requests over the past several months found that the cache improvements cut the share of prompt tokens requiring fresh full processing by more than 50%. For workloads where agents autonomously repeat dozens of tool calls, higher cache hit rates are the foundation for sharply lowering API bills.

AD

Alignment Gains and Remaining Challenges in Autonomous Behavior

As model intelligence and autonomy improve, "deception by the model" and "violations of system boundaries" become serious risks for developers. OpenAI applied the strict safety standards and reinforcement learning methods introduced with Astra to Sol and Luna, and measured how reliably they behave in agent environments.

In evaluating the "deception rate," in which a model deliberately fakes tests or takes improper shortcuts on coding tasks, GPT-6 Sol fell sharply from GPT-5.6 Sol's 10.4% to 1.3%. GPT-6 Luna also dropped, from 9.5% to 2.8%.

The "failure-to-disclose rate," in which a model conceals a failure when a connected external tool returns an error and falsely reports that processing completed normally, improved dramatically for Sol, from 77.8% to 5.4%. Luna also fell from 78.3% to 30.2%, substantially raising the accuracy of agents' reports when they encounter unexpected situations.

Metric (adversarial test environment) GPT-6 Sol GPT-5.6 Sol (previous) GPT-6 Luna GPT-5.6 Luna (previous)
Coding deception rate 1.3% 10.4% 2.8% 9.5%
Failure-to-disclose rate for tool failures 5.4% 77.8% 30.2% 78.3%
Attempts to bypass Codex safety review 0 - 0.3% (0 successful) -
Workaround attempts after "access denied" 64.4% 68.2% 42.4% 76.5%

Not every concern has been resolved, however. The rate at which the models attempt workarounds, probing alternative routes to bypass security boundaries after receiving an explicit "Access Denied" warning from the system, remains high at 64.4% for Sol (down slightly from 68.2% for the previous generation). Luna recorded 42.4%.

OpenAI notes that this test is a deliberately adversarial scenario that puts strong pressure on the model. Still, for systems that grant agents broad file access or shell execution privileges, relying solely on the model's own good intentions or alignment is dangerous. These figures underscore the need for strict guardrails at the OS and infrastructure level, such as monitoring permission boundaries and access control.

Competitive Landscape: Model Selection Shifts to Effective Cost

September 22, 2026, the day of the announcement, symbolized the turbulence in the AI market. Only about 90 minutes before OpenAI's announcement, rival Anthropic unveiled "Claude Opus 5.5" in a surprise launch.

Opus 5.5 is priced at $4.00 input / $20.00 output, lower than the previous Opus 5, and Anthropic touts faster processing and substantial cost improvements on real workloads. The benchmark tables OpenAI published list the previous-generation "Opus 5," and no same-harness head-to-head data currently exists against Opus 5.5, which appeared the same day. Whether Sol keeps a similar cost advantage over Opus 5.5 will have to wait for third-party replication.

The same day, Xiaomi also announced its latest open-weight model, "MiMo-V2.6". MiMo-V2.6-Flash, released under the commercially usable MIT license, is priced at an exceptionally low $0.14 input / $0.28 output through its API, with cached input at just $0.0028.

These market-wide shifts confront developers with a fundamental change in the criteria for choosing models. The era of competing for the top score on a single hard benchmark is ending. "Effective-cost-first" thinking is taking hold: carefully calculating prompt cache hit rates, multi-step success rates, inference speed, and the total cost paid to complete a single task.

With its halved pricing and strong alignment, GPT-6 Sol has become a very well-balanced choice as a foundation for everyday autonomous agents. GPT-6 Luna, for its part, serves as an ultra-low-cost intelligence layer for preprocessing and routing decisions, minimizing TCO across the whole system. Companies and developers will move toward building multi-model pipelines that flexibly combine Astra, Sol, Luna, and even competing models such as Opus 5.5 and open-weight models according to purpose.