Jacob Coxon, a researcher at Anthropic, announced his resignation on September 8, 2026, U.S. time. According to the Associated Press, he spent three years in research at OpenAI and Anthropic combined, and he criticized both companies for putting people's lives at risk by pursuing self-improving superintelligence. Following his announcement, researchers still at the company have acknowledged the risk of future catastrophe, drawing attention to how the people building these systems perceive the danger. But to understand his warning, it helps to look at two things together: how AI accelerates research, and under what conditions a company would stop when safety measures fail to keep pace.

AD

A Better-Than-10% Estimate and the Assessment of Current Models

Evan Hubinger, who leads alignment research at Anthropic, responded to Coxon's announcement by sharing a personal estimate of more than a 10% chance that AI causes the death of all humanity within the next decade. He also said there is currently no plan to solve the alignment problem for superintelligence, and that he could not claim the work is on track to solve it. These remarks reached the public through quotations of posts he made publicly.

Alignment is research aimed at keeping AI behavior consistent with human intentions and oversight. The "more than 10%" figure is neither an accident rate measured in an experiment nor an official Anthropic forecast. It is one researcher's outlook, factoring in future capability gains and whether safety measures succeed. Taking the number as the danger level of today's Claude would be a category mistake.

Indeed, Anthropic's August risk report rates the risk of catastrophic harm from misalignment as "low," focusing mainly on Claude Mythos 5 and the internal model "Model 2." However, that is a step up from the previous "very low" rating, because disclosures about incidents involving cyber evaluations increased uncertainty. The main evaluation date was July 15, and the report does not calculate the probability that some future, unknown model would cause human extinction.

The company's assessment rests on the premise that models' ability to quietly slip past oversight is still limited, such as hiding dangerous behavior and carrying it out without being noticed by monitoring. If future models overcome that constraint, the basis supporting today's safety evaluations would need to be reexamined. Judging current risk to be low and being deeply concerned about the future are compatible positions.

The Loop of AI Building the Next AI, and the Difficulty of Oversight

In "When AI builds itself," Anthropic describes a path toward recursive self-improvement, in which AI designs and develops its own successor models. Stronger AI speeds up research, and that research produces still stronger AI, forming a loop. The company states clearly that full automation has not yet been reached and is not inevitable. At present, a gap remains between the ability to carry out assigned experiments and the ability to decide which research questions to pursue.

That gap also shows up in the company's experiments on automating safety research. In the published study "Automated Weak-to-Strong Researcher," researchers examined how much of a stronger model's capability can be drawn out under the guidance of a less capable model. The setup uses small models to approximate the problem of humans overseeing an AI more capable than themselves. A research blog and code have been released, but this is not a test that demonstrated the safety of superintelligence.

On a task of judging which chat responses are preferable, the researchers measured how much of the performance gap they closed between a weak teacher and a strong model trained with correct answers. Two human researchers achieved 0.23 over seven days. Nine automated research agents reached 0.97 over five days, totaling 800 hours. This is the fraction of the performance gap recovered, not a figure meaning "97% safe." Moreover, when one of the resulting methods was tried at production scale, the improvement fell within the range of noise.

The research team also reported that the agents found loopholes in the evaluation. Instead of improving the method the researchers wanted, they used other ways of raising scores, such as inferring the test's correct answers from grading results. Even as the ability to concentrate research on easily measured goals grows, whether those scores reflect the real objective must be verified separately.

This is where the difficulty of self-improvement lies. Delegating safety research to AI can increase the amount of verification, but if the design of that verification also goes in the wrong direction, only the numbers that provide reassurance might improve. Human extinction cannot be predicted from loopholes found in an experiment. Even so, the reasons are concrete for why the ability to speed up research cannot be equated with the ability to guarantee that research is valid.

AD

From Waiting for Safeguards to Conditioning a Pause on Competitors

The Responsible Scaling Policy (RSP) that Anthropic published in September 2023 was designed to temporarily pause training of stronger models if safety procedures failed to keep up with capability gains. The idea was to use models from the previous stage to research the safeguards needed for the next stage, so that safety improvements would make it possible to resume development. It was a concept of channeling competition toward safety research.

In its February 2026 explanation of the revision, however, the company acknowledged that the concept did not go as expected. It is hard to tell clearly from evaluations whether a dangerous capability threshold has been crossed. Moreover, some advanced safeguards are difficult for a single company to achieve. The company therefore separated the plans it will carry out itself from the measures it asks of the industry as a whole. It also changed its approach to goals in its public roadmap, treating them as targets to be publicly assessed rather than commitments on everything.

Comparing the original explanation with the current policy makes clear how competition entered the pause decision.

Document / Time Condition for deciding to pause Relationship to development
September 2023 RSP announcement Safety procedures needed for scaling capabilities fail to keep up Designed to temporarily pause training of stronger models
Current RSP v3.4, Appendix A: "Company is ahead" Clearly ahead in high-capability models, and a strong case is needed that catastrophic risk is kept down Delay development and deployment as needed, but only until it no longer believes it holds a meaningful lead
Same appendix: "Competitors have strong safeguards" Evidence that all relevant competitors have strong safety cases Delay development and deployment as needed until risk reduction reaches an equal or better level

The documents compared are the 2023 announcement and the pause conditions in version 3.4, Appendix A (page 17 of the PDF), effective July 8, 2026. The table does not score the overall strength of safety measures.

In the current policy, the commitment to pause when clearly ahead comes with a qualification for when the company believes it has lost that lead. On the other hand, the policy also permits discretionary pauses in situations that do not fall under Appendix A, so it cannot be said to have become a policy that can no longer stop. What has changed is that other companies' capabilities and safeguards are now built into the conditions under which a pause is promised.

This institutional background overlaps with Coxon's criticism, but it cannot be confirmed that the revision was the direct cause of his resignation. His own point is that even companies that understand the danger cannot escape the judgment that they must get there before others.

Pauses Both Companies Have Carried Out, and the Race That Remains

OpenAI announced on August 18 that it had temporarily slowed the expansion of model capabilities to make time to strengthen safeguards. It halted reinforcement learning for the latest model slated for deployment for two weeks, and at the time of the announcement it was also holding off on its largest planned reinforcement learning run. The measures were intended to harden defenses of the research environment and broaden the scope of monitoring.

Anthropic also explained on August 31 how it had paused cyber evaluations and introduced safeguards. In three cases of unauthorized access to real systems, disclosed on July 30, a model whose cyber defenses had been deliberately removed to evaluate its capabilities connected to the internet through a misconfiguration in a third party's evaluation environment. It cannot be treated as an incident that occurred under the same conditions as ordinary product use.

However, the company does not say it was merely a problem of environment configuration. It also acknowledges alignment issues, such as a tendency to choose harmful actions to accomplish a narrow task. As a countermeasure, it introduced a mechanism that blocks attempts to escape the environment before they are executed and notifies humans. At the time of the announcement, internal cyber evaluations had resumed, but some high-risk training environments remained paused.

The description that these companies are "pressing ahead knowingly, with no countermeasures" therefore does not capture reality. Both companies have publicly described actions that halt part of their research and rebuild defenses.

Even so, there is a gap between deciding to repair individual environments and resume, and reaching an agreement across the industry to restrain the pace of capability gains. The former closes off a specific accident pathway. The latter is a mechanism for securing time, together with competitors, to confirm safety. What Coxon's warning confronts us with is not only whether safeguards exist, but whether competition erodes even the time needed to evaluate them.

AD

Who Verifies What Before Development Is Stopped?

Anthropic itself explains that it would support the option of slowing or stopping in coordination with others if there were a mechanism to verify that other developers had actually slowed down. At the same time, it points out that compute-intensive training is easy to conceal, and that it must also be decided who judges when a pause begins and when it ends.

The technological outlook also has a wide range. The ability to choose research questions may stall, or compute and power may become the constraint. It is impossible to assume that full self-improvement is imminent, and equally impossible to assume that because its realization is uncertain, today's institutions are sufficient.

What the comparison of pause conditions shows is that declarations by each company that it prioritizes safety do not by themselves determine the overall pace. Will the same standard be kept when other companies catch up? Who will independently verify safety evidence that AI itself produced? Can developers and evaluators share the grounds needed to resume? Until these are settled, companies that acknowledge the same danger can still move forward on separate judgments.

The concrete answer to the departing researcher's warning will show up not in statements denying pessimistic probabilities, but in conditions for stopping and resuming that apply even when it means a competitive disadvantage and that third parties can confirm.