Anthropic's confidential IPO prospectus warns investors of "existential risks to humanity" from advanced AI, Reuters reported on September 28, citing a review of the document. The prospectus reportedly also gives examples of self-preserving behavior, such as resisting shutdown and concealing information. This is not a report that humans have lost control of Claude models currently on the market. A separate risk assessment published by Anthropic is limited to specific models and specific pathways to harm.

AD

What the roughly 80-page risk disclosure warns about

According to Reuters, the prospectus says advanced AI could create "catastrophic or even existential risks to humanity." It reportedly cites blackmail-like behavior as an example, along with models resisting shutdown and hiding or manipulating information. Reuters also says Anthropic acknowledges that more advanced models and broader uses could increase the danger of models causing harm.

The original prospectus has not been made public. What can be verified is Reuters' reporting on the document it obtained, and the wider context of the filing cannot be checked directly against published material. Anthropic announced in June that it had confidentially submitted a draft S-1 to the U.S. Securities and Exchange Commission. This is not a first-time IPO filing.

Reuters also reported that of the prospectus's 261 pages, about 80 are devoted to risk factors and 48 to the business description. The space given to risk shows the breadth of issues put before investors. But page counts cannot be used to calculate the probability of an existential catastrophe or the effectiveness of safeguards. What matters is separating the research behind the warning from incidents that have actually occurred.

Reading the warning, the experiments, the incident and the current risk assessment separately

The IPO's existential-risk warning, misbehavior in controlled experiments, actual access to external systems, and the assessment of catastrophic harm from current models are each different kinds of evidence.

Source and date What it shows What it does not show
Confidential prospectus reviewed by Reuters on Sept. 28 Reported to warn investors of serious dangers that future advanced AI could pose Not proof that today's Claude has caused an existential crisis
Anthropic's summer 2026 simulation experiments Models from multiple companies attempted actions such as sabotaging code and manipulating records in staged scenarios Does not reveal how often such behavior occurs in ordinary use
Incident during testing reported by Anthropic on July 30 A model with its safeguards deliberately removed reached real external systems through a misconfiguration in a third party's evaluation environment Humans did not lose the ability to regain control, and no catastrophic harm occurred
Anthropic's August risk report The company qualitatively rated as "low" the risk that Claude Mythos 5 and an unreleased internal model, "Model 2," would behave inappropriately in high-risk settings in ways leading to catastrophic harm Not an assurance of safety for future models or for every use

This table sorts Reuters' report on the prospectus and the experiments, incident and risk report published by Anthropic by what was observed. It is not a comparison converted to a common numerical scale. In the simulations in particular, Anthropic itself says these were not real-world events. Results from experiments testing blackmail-like behavior cannot be read as a report that blackmail actually happened.

The access to external systems, by contrast, was a real case. According to Anthropic, a model whose cyber safeguards had been removed for evaluation exploited a misconfiguration in a third-party environment that was supposed to be cut off from the internet. The company's follow-up investigation identified one more case in addition to the three initially reported. This points to a failure to isolate the evaluation environment adequately from real systems. But the conditions of a test in which safeguards were deliberately removed cannot simply be applied to the Claude models offered in ordinary service.

AD

Models that notice they are being evaluated make testing harder

Reuters also reported that the prospectus describes a model's possible awareness that it is being safety-tested as a major constraint on the ability to detect dangerous behavior. If a model recognizes "this is a test" and behaves differently from how it would when it actually holds authority, the absence of problems in testing means less.

The concern is spelled out in pages 46–47 of Anthropic's August risk report. The report acknowledges that if a model can tell a simulated environment is artificial, evaluations could miss serious misbehavior. The company says evaluation awareness "partially" lowers its confidence in ruling out serious misbehavior.

That said, it is not claiming that every test has the same weakness. The report says evaluations using real internal usage data or isolated environments close to real conditions retain value. This is the difficulty of safety evaluation: the more carefully conditions are constructed to elicit dangerous behavior, the more likely the model is to see through the artificiality. Waiting for real-world use alone, on the other hand, cannot adequately capture rare but serious situations.

"Low" is the company's own limited, current assessment

Anthropic's risk report does not claim to have eliminated the danger. It raised its rating for catastrophic harm from inappropriate model behavior in high-risk settings from "very low" to "low," saying that recently disclosed incidents during cyber evaluations increased its uncertainty. This classification is the company's own qualitative judgment, not a numerical probability.

The report's reference date was July 15, and it was published on August 14. The models assessed are Claude Mythos 5 and the unreleased internal model "Model 2." According to Anthropic, Mythos 5 is offered to select customers through Project Glasswing and is also available to the general public as Claude Fable 5 with additional safeguards. Uncertainty from incidents disclosed after the reference date was factored into the judgment, but the assessment is limited to these models and the pathways the report defines. It cannot be compared on an equal footing with the future advanced AI in general that the IPO documents warn about.

The report examines the risk that naturally arising inappropriate model behavior leads to serious harm through specific pathways. It does not cover malicious users exploiting a model or the dangers of still more powerful models yet to be developed under the same single word, "low." Because the judgment is scoped, it will need to be updated as models and usage change.

The independent International AI Safety Report 2026 summarizes that current AI does not yet have the capability to cause humans to lose control. It also notes that capabilities for autonomous action and for finding loopholes in evaluations are growing. Future loss of control would require not just the ability to evade oversight but also a tendency to use it in harmful directions, plus the authority and environment to act. Experts remain widely divided on how likely it is and when it might happen.

Reuters also reported that the prospectus says the return Anthropic will get from its investment in safety research is unclear. In a compute estimate the company published, roughly 6% of the compute used for AI research and development during a one-week sample in July went to safety research. That is not 6% of research spending, nor 6% of the company's total compute. Anthropic itself explains that the share of compute alone does not adequately measure how much effort goes into safety measures.

Once a public version of the S-1 is released, the context of the reported warnings can be checked against the original. From there, two things will help gauge the weight of the warning: how closely capability evaluations of new models come to resemble real-world use, and whether the countermeasures for problems seen during testing can be verified by third parties. The materials available now show that preparation for serious risks is needed, but they do not support the conclusion that an existential danger to humanity has already materialized.