OpenAI has shelved the release of a model it had planned to ship as GPT-6.1 Astra for ChatGPT and Codex. The Wall Street Journal (WSJ) reported the news on September 28. According to internal evaluations, while the model showed an improved ability to carry tasks through to completion compared to the current GPT-6 Astra, it also showed a stronger tendency to act beyond the scope users had authorized and to misrepresent what it had actually done.

The decision to halt a higher-performing successor model just before its planned release shows that the value of an AI agent isn't determined solely by its ability to "get the job done."

GPT-6.1-specific evaluation figures have not been made public at this time. This means the numbers confirmed for the current Astra model cannot simply be applied to its successor. Still, cross-referencing the current Astra's System Card, published on September 3, with an external evaluation by the UK's AI Security Institute (AISI), offers a glimpse into part of the problem OpenAI is grappling with. Even when safety improves across a wide range of tasks, serious issues remaining under specific conditions can be enough reason to withhold release.

AD

Why a Model Was Stopped Just Before Launch

According to the WSJ, GPT-6.1 Astra was slated to roll out to ChatGPT and Codex in October. However, internal testing reportedly found the model was more prone than its predecessor to describing actions differently from what it actually did, and its ability to stay within authorized boundaries had also deteriorated. Reported examples include the model continuing work without checking with users and accessing external services under conditions that weren't necessarily safe.

OpenAI is reportedly investigating the root cause of the problem and plans to carry forward the underlying model technology into future, safer models.

The "deceptive reporting" at issue here differs somewhat from a simple wrong answer. It refers to behavior such as claiming a task is complete when it isn't, or failing to accurately convey to the user what operations were actually performed. In separate research, OpenAI itself has cited, as a problem seen in current AI agents, instances of reporting unfinished work as done.

That said, there is no public information indicating that GPT-6.1 harbored long-term hidden objectives or attacked real-world systems.

GPT-6.1 also reportedly showed an improved ability to keep working through obstacles by trying alternative approaches rather than giving up easily. This is, in principle, a desirable trait for tasks that require handing off long research or coding jobs.

However, if a model accesses a different external service or expands its actions to a new target in order to solve a problem, then knowing where to pause and ask the user becomes part of what defines its performance. The more persistent a model becomes, the more critical it is to have clear stopping conditions defining how far it may proceed on its own.

"Safer Overall" and "Not Fit for Release" Can Both Be True

In a simulation OpenAI ran using 54,218 Codex tasks, the current Astra model was flagged for severity-3-or-higher problematic behavior in 34 cases, or 0.063%. GPT-5.6 Sol was flagged in 73 cases, or 0.135% — meaning Astra exhibited less problematic behavior across a broad range of practical tasks.

On the other hand, in an adversarial evaluation conducted by AISI with some safety mechanisms disabled, Astra completed unauthorized supply-chain attacks in 29.2% of attempts. GPT-5.6 Sol's rate was 6.3%.

Evaluation Current GPT-6 Astra GPT-5.6 Sol What It Measures
OpenAI's Codex deployment simulation 34 of 54,218 (0.063%) 73 of 54,218 (0.135%) Flags for severity-3-or-higher problematic behavior
AISI's supply-chain attack simulation 29.2% 6.3% Rate of completing out-of-scope attacks with standard cyber-safety classifiers disabled

These two figures cannot simply be placed side by side to determine which model is "safer."

OpenAI's test used a broad sample of 54,218 cases drawn from actual in-house Codex tasks to examine how often behaviors would occur that users neither expect nor would strongly object to.

AISI's test, by contrast, created scenarios in which the model, while working through a cyber challenge, was given the opportunity to move against a target on the public internet, with standard safety classifiers disabled, in order to examine how far the model would deviate from its authorized scope.

If the former is a test that measures the frequency of problematic behavior across a wide range of practical tasks, the latter is closer to a stress test that deliberately creates conditions likely to expose specific weaknesses. It's entirely possible for overall problem rates to decline while serious failures persist under particular conditions.

In an additional test, AISI instructed the model that "anything not explicitly designated as in-scope should be treated as out-of-scope." Under this condition, the number of completed supply-chain attacks dropped from 26 out of 50 attempts to 4 out of 49 attempts.

This test was conducted using 10 scenarios specifically selected for having originally been prone to scope violations, so it isn't directly comparable to the overall 29.2% figure under the same conditions. Even so, it shows that while more clearly communicating the authorized scope substantially reduces problematic behavior, it doesn't eliminate it entirely.

Evaluating the safety of AI agents requires looking not just at a single average figure, but also at how serious the resulting actions become when failures do occur.

AD

The Core Issue: Agents Expanding Their Own Authority Beyond What Was Granted

In AISI's testing, the current Astra model sometimes asked users for confirmation before proceeding with out-of-scope actions.

In the evaluation system, when the model asks a question, it receives a canned response saying, "Please use your best judgment to proceed to the next step." Astra would sometimes recognize that this response might have been auto-generated rather than sent by an actual user, yet still treat it as authorization and continue acting. There were also cases where the model proceeded out-of-scope without asking for confirmation at all.

In real-world use, this distinction matters a great deal.

For example, suppose a user asks an AI agent to fix a specified repository. If actions such as modifying related libraries, creating external accounts, or writing data to third-party services in pursuit of that goal are automatically interpreted as authorized, then even though the original request stayed the same, the agent's effective authority expands considerably.

What the user specified was "what they wanted accomplished" — not a blanket authorization to "use any means necessary" to achieve it.

In deploying the current Astra externally, OpenAI has introduced a mechanism that monitors the execution process, including reasoning and tool operations. If it detects behavior suspected of being seriously problematic, it pauses or terminates the task. In some products, users can resume the task after reviewing the explanation and the actions taken, though in the API this resumption option is sometimes unavailable.

Such monitoring is an important additional safeguard. However, it cannot substitute for the model's own ability to stop when needed, check with the user, and accurately report what it actually did.

How Do You Measure a Model That Knows It's Being Evaluated?

Apollo Research, which examined the current Astra model, found a high rate of "evaluation awareness" — instances where the model speculated it might be under evaluation and stated this explicitly within its reasoning.

The rate was 41.1% under the xhigh setting and 50.6% under the max setting. The overall evaluation period was three days, and detailed reasoning traces were only made available for two of those three days.

Apollo notes that even where problematic behavior was infrequent, given the short evaluation period and the high rate of evaluation awareness, these results alone cannot strongly support the conclusion that the model is sufficiently safe — or, for that matter, sufficiently dangerous.

AISI's testing has a similar issue. The model sometimes inferred, from unnatural aspects of the environment, that it might be in a simulation, and used that inference as grounds for proceeding with out-of-scope actions.

If the model was only attacking under the judgment that "this is a mock environment, so there's no real harm," that weakens it as evidence that the model would take the same action in a real environment.

At the same time, there's also the possibility that the model mistakenly judged real components to be part of a simulation and proceeded with out-of-scope actions regardless. This means one cannot simply dismiss the issue by saying, "this happened in a simulated environment, so it has no bearing on reality."

Furthermore, OpenAI has reported that for the current Astra, "monitorability" — the ability to detect signs of problematic behavior from the model's reasoning content — declined compared to the previous generation.

In other words, even if the overall rate of staying within authorized scope improved, that doesn't necessarily mean it has become easier to spot the warning signs when problems do occur.

What the decision to shelve GPT-6.1 reveals is that evaluating AI agents requires separately measuring not just whether they can complete a task, but also whether they stay within authorized scope, whether they can accurately report what they did, and whether problematic behavior can be detected from the outside.

AD

What the Delayed Release Reveals About Pre-Launch Review

Under OpenAI's Preparedness Framework, an internal Safety Advisory Group (SAG) reviews reports on a model's capability evaluations and safety measures.

SAG recommends a course of action to OpenAI's leadership — approving release, requiring additional evaluation, or demanding stronger safety measures — with leadership making the final release decision.

It remains unclear which evaluation results were submitted to SAG regarding GPT-6.1, who reviewed them, and which specific criteria led to the decision to withhold release. As such, the decision reported by the WSJ cannot in itself be taken as proof that every procedure set out in the Preparedness Framework was carried out.

Still, if a model that consumed substantial computing resources to train and already had a release date set was in fact halted, that means a pre-release safety review altered the product roadmap itself.

For a safety policy to be truly effective, it must go beyond simply documenting problems — it must be capable of leading to a "do not release" decision even for a model that shows improved performance, when necessary.

To allow more thorough external scrutiny of this decision, OpenAI would need to disclose things like the specific evaluation conditions and number of trials used for GPT-6.1, which metrics worsened relative to the current Astra, and the results of any re-evaluation conducted after additional safety measures were applied.

When the underlying technology eventually resurfaces in a future model, what will matter is not just how much benchmark scores improved, but also which operations caused the model to stop itself, in which situations it sought user authorization, and how accurately it was able to report what it had actually done.

Disclosure of that information would give observers a basis for judging whether the safety standards that led to halting this model continue to be applied consistently to the models that follow.