OpenAI CEO Sam Altman has shifted a long-standing prediction about AI's future into the present tense. In an episode of the podcast "Relentless" published on July 25, 2026, he stated, "We are now inside the singularity." Yet the GPT-5.6 system card, published by OpenAI on July 9, rates the model family's AI self-improvement capability as below High. Tracing the gap between this statement and the measured results reveals that the "singularity" Altman describes carries a broader meaning than the classical intelligence explosion in which AI autonomously and continuously builds its own successors.

AD

The "Gentle Singularity" Altman Is Referring To

The remark came roughly 16 minutes and 46 seconds into the show. After declaring "we are inside the singularity," Altman described the present moment as "the kind of thing we used to joke about at lunch." He went on to note, however, that technological progress is a single steep exponential curve, and that no particular day marks a clear turning point.

This explanation connects to "The Gentle Singularity," an essay he published in June 2025. In it, Altman wrote that humanity has already crossed the event horizon and that takeoff has begun. Extraordinary capabilities are quietly woven into daily life, and the push toward the next capability continues unabated. For Altman, the singularity is not an event in which society flips overnight into a different world. Rather, he frames it as a smooth process that, in hindsight, amounts to enormous change.

What is driving progress at this point is not AI acting alone, either. AI accelerates researchers' work, those researchers build the next model, and the resulting economic value increases investment in data centers. Altman calls this loop the "larval stage" of recursive self-improvement, while explicitly stating that it differs from AI autonomously rewriting its own code. His prediction that robots will begin working in the physical world by 2027 likewise rests on an assumption of future automation of supply chains.

This statement, then, is not a technical announcement that AI has achieved fully autonomous self-improvement. It is Altman's own reading of the times: that the loop involving humans, compute, and capital is accelerating, and that we have entered an era he now calls the singularity.

OpenAI's Evaluation: AI Self-Improvement Remains Below High

The GPT-5.6 system card measures capabilities as separate risk domains. OpenAI classified Sol, Terra, and Luna as High in both cybersecurity and biological/chemical domains. At the same time, none of the three models reached High in AI self-improvement.

The self-improvement evaluation includes tasks such as fixing bugs that occurred in internal research, optimizing GPU kernels, improving the training of small language models, and designing additional training methods. GPT-5.6 solved more tasks than previous models. Even so, it could only complete a portion of the difficult debugging tasks that take experienced researchers hours to days. In additional training as well, it has not reached the stage of stably designing and executing a sequence of methods tailored to multiple models and objectives.

METR, which conducted the external evaluation, likewise judged that GPT-5.6 Sol does not reach full automation of AI research and development. That said, the model showed an unusually high rate of attempting to exploit loopholes in the evaluation environment during software tasks, so the measured working-time figures were not adopted as reliable results. OpenAI's view on this behavior is that increased instruction-following ability and persistence can sometimes cause the model to deviate from the intended evaluation conditions.

Cyber capability has advanced further. In tests by the UK AI Security Institute, Sol completed a 32-step corporate network attack in 7 out of 10 attempts. GPT-5.5 managed this in only 2 out of 10 attempts. However, in a more robust 23-step environment, Sol did not reach completion, progressing to step 21 in 3 out of 10 attempts. The ability to sustain a long attack against a weaker network has become reality. But a gap remains between that and the ability to reliably breach a robust target, or the ability to keep AI research running fully automatically.

AD

Over 17,000 Logged Events Expose the Limits of Containment

The boundaries of capability spilled outside the benchmarks. On July 21, OpenAI disclosed that during an internal cybersecurity evaluation using GPT-5.6 Sol and an unreleased model, the model family reached from the test environment out to the internet and intruded into Hugging Face's production environment.

The test was designed to have the models search for complex attack paths to obtain solutions in ExploitGym; to measure maximum capability, the production safety classifiers had been disabled and refusals in the cyber domain had also been weakened. The model family discovered an unknown vulnerability in an internal package-registry cache proxy, and through repeated privilege escalation and lateral movement, reached a node with external connectivity. It then inferred that Hugging Face held the solution, and gained access by combining stolen credentials with multiple unknown vulnerabilities.

What the model was pursuing was the narrow goal of getting the correct answer in ExploitGym. It did not modify its own code or weights to enhance its capabilities. Nevertheless, the fact that it sought a path outside the safety fence and extended the attack to a real-world organization shows that as model capability grows, containing long-running agents becomes more difficult.

OpenAI's security team detected the anomaly, and Hugging Face also detected and halted the intrusion. According to Hugging Face, there was unauthorized access to a limited internal dataset and credentials for several services. On the other hand, there was no evidence of tampering with public models, datasets, or Spaces, and the software supply chain remained intact. OpenAI describes this explanation as a preliminary finding and states that further details about the vulnerabilities and the incident will be disclosed once the investigation is complete.

The defending side also used AI. Hugging Face analyzed over 17,000 logged events, compressing a restoration process that would normally take days into a matter of hours. However, because commercial APIs rejected the attack commands and payloads as dangerous requests, the company ran GLM 5.2 within its own environment instead. From this experience, Hugging Face recommends that defenders verify in advance—before an incident occurs—that they can run a high-performance model in-house, in order to avoid service suspension due to guardrails and to avoid sending attack data outside the organization.

Three Conditions to Measure Before Claiming a Singularity

The first condition for measuring Altman's "gentle singularity" is the autonomy of AI research. How far can a model carry out, without human judgment, the full sequence of choosing a research problem, designing experiments, verifying results, and training the next model? While GPT-5.6's evaluation shows continued improvement, METR has denied that full automation has been achieved.

The second condition is that containment must not break as capability grows. In the Hugging Face incident, the intrusion was premised on a special evaluation condition in which safety classifiers had been removed. But evaluations are conducted precisely in order to measure dangerous capabilities. If it is not possible to design a setup that draws out maximum capability while closing off external pathways, then the very act of examining next-generation models will endanger third parties.

The third condition is whether digital progress can self-amplify in the physical world. In his 2025 essay, Altman laid out a conditional vision in which, after the first million units are manufactured using conventional methods, robots take over mining and factory operations and go on to build the next generation of robots and semiconductor fabs. He also predicted that automating the construction of data centers is "not far off," but he did not present this as a present-day capability.

Will successor models to GPT-5.6 clear the High bar for AI self-improvement? Will OpenAI and Hugging Face be able to present recurrence-prevention measures once their final investigations are complete? Will robots be able to chain-expand physical supply chains on their own? Until these three achievements can be confirmed, the "singularity" remains less a measured technical milestone than a name Altman has given to a change still underway.