On September 12, OpenAI CEO Sam Altman said the company would accept independent evaluators with access close to that of its own employees. The statement responded to the proposal by Anthropic CEO Dario Amodei to slow the pace of AI development. Elon Musk and Demis Hassabis also voiced support the same day. But the three did not commit to the same oversight system. A close reading of the original posts shows a gap between a concrete pledge to bring evaluators inside a company and support for the general direction of a proposal. Challenges also remain in turning this into a mechanism shared across the industry.

AD

The three endorsements are not the same commitment

In a post on X, Altman said he agreed with Amodei's view that progress in frontier AI needs to be paced, and revealed that it had been one of the main topics at OpenAI in recent weeks. He also praised the idea of giving independent evaluators access close to that of employees, and said OpenAI would work toward the same. He said details would be shared soon.

This goes further than a routine compliment for a rival's proposal, because Altman described what OpenAI itself would do. Musk's post, by contrast, was a brief statement of agreement: "Dario is right." It said nothing about how evaluators would be accepted or how far development should be slowed.

Hassabis said the essay points in the right direction, adding that the details still need to be worked out. He then tied this back to his earlier proposal, noting that this is why he had himself suggested an industry-wide standards body for frontier AI.

Sorting what the September 12 posts explicitly stated into commitments to act and reservations makes the differences clear.

Speaker Position stated in the post Commitment or reservation on accepting evaluators
Sam Altman Supports pacing development and independent evaluators with employee-like access Says OpenAI will work toward the same; details to come
Elon Musk Brief statement of agreement with Amodei No mention of specific measures
Demis Hassabis Supports the direction; refers to his own industry standards body proposal Details need work; does not state acceptance of embedded evaluators

Of the posts that day, only Altman's went as far as stating that evaluators would be accepted. The posts by Musk and Hassabis did not record the same pledge.

This comparison lines up what the posts said; it is not a comparison of measures each company has already implemented. Altman's statement also did not announce that evaluators had begun work or that contracts had been signed. Broad support should be read separately from the establishment of a unified slowdown rule.

Evaluators inside the company, standards across the industry

Amodei's essay, "We Must Pace the Frontier", proposes a mechanism in which external evaluators continuously verify safety measures from inside a company. Beyond testing finished models, they would check whether the training process and safety commitments are actually being honored. Anthropic pledged to accept this arrangement on its own and called on other companies to follow.

Hassabis's proposal, published July 14, centers instead on creating evaluation procedures that can be used across companies. A standards body would be established, led by the United States, either as a public-private partnership under federal oversight or as a self-regulatory organization. It would set capability-evaluation thresholds, and models meeting them would be treated as frontier-class.

In short, resident evaluators would see inside each company's development operations, while a standards body would align which tests each company is evaluated against. If the two work together, they could potentially link pre-release model testing with verification of the processes leading up to it. Amodei himself mentions Hassabis's proposal as an example of consultation through an industry body with ties to the government.

Hassabis's proposal also includes a path from voluntary participation to formal requirements. Initially, up to 30 days before release, companies would voluntarily submit models for review. Once the evaluation procedures proved effective, the plan envisions moving to a system requiring frontier models deployed in the U.S. market to pass. There is currently no submission requirement, nor any demand for a 30-day halt to development.

Another notable point is that coverage would not be determined by a company's size or name. Frontier-class models could be covered whether open or closed, and regardless of the country in which they were developed, while models that fall short of that level would be excluded. It is not a plan to exempt startups or universities across the board simply because of who built the model.

The tests themselves would not be fixed, either. They would initially be reviewed on a schedule such as quarterly, with outdated evaluations replaced. In the future, the standards body could create confidential tests independent of companies, to prevent models from being optimized only for known problems. If the tests that measure danger are left behind by model progress, only a certificate of passing would remain.

AD

Independence rests on access and the right to publish

A notable condition in what Anthropic has pledged is that evaluators can publish unfavorable findings. The essay says contracts will include the right to publish, free from Anthropic's editorial control, not only significant findings about risks and incidents but also which information the evaluators could and could not access.

If evaluators can see internal information but cannot release conclusions inconvenient to the company, third-party verification is limited. Conversely, if access to information is insufficient, a right to publish alone cannot confirm the substance. The Amodei proposal tries to secure both by granting authority roughly equivalent to that of internal risk-assessment staff and allowing evaluators to gather information through conversations with employees.

The access is not unlimited, of course. Constraints under law and contracts, and protection of customer and partner information, remain. Safety-sensitive information and legally protected information can be redacted in limited ways, but findings cannot be hidden merely because they are unfavorable. If redactions materially affect a conclusion, evaluators could point that out publicly. The idea is that even the exceptions made to protect secrecy can be scrutinized from outside.

Under Hassabis's proposal, who runs the standards body is also an issue. He outlined a plan to include independent technical experts and open-source representatives on the board, with industry primarily funding the talent and computing resources needed. The design challenge is how to absorb knowledge from the front lines of development while remaining independent of the companies providing the funds.

The U.S. Financial Industry Regulatory Authority (FINRA), cited as a reference example, is not a body where companies merely exchange views voluntarily. According to FINRA's own description, it is a private nonprofit funded by membership fees and overseen by the U.S. Securities and Exchange Commission (SEC). Under federal law, it supervises its member securities firms, writes and enforces rules, conducts examinations, and has the power to expel violators from membership where appropriate.

If this precedent were transferred to AI, gathering experts alone would not complete the system. It would be necessary to decide which agency supervises and what can be done to companies that do not comply with evaluations. The role of resident evaluators in finding and publishing problems is separate from the authority to halt development or deployment. The statements of support so far have not settled that allocation of authority.

From agreeing on a slowdown to verifiable commitments

Behind the case for slowing down is concern about "recursive self-improvement," in which AI helps develop the next AI and those results further accelerate development. However, Anthropic itself has not said that a state in which AI fully autonomously designs and develops its successors has been achieved. Its "When AI builds itself" states clearly that we have not yet reached that point and that it is not inevitable.

The automation of development support now under way should be distinguished from a future fully autonomous loop. Even so, the proposal that third parties continuously check whether safety measures are keeping up with growing capabilities holds up. Amodei also explains that pacing does not mean stopping model training or technological progress.

Coordination among companies will not advance on the number of endorsements alone, either. Amodei notes that discussions on safety raise antitrust issues, and calls for government mediation and limited exemptions for some consultations. He proposes building common standards among companies in democratic countries and then extending them to international cooperation, including with China, but he himself is cautious about achieving a full slowdown worldwide.

In OpenAI's further explanation, who the evaluators are, when they begin, and what information they can access will be concrete factors for judgment. It is also not yet clear from the post whether OpenAI will adopt a condition equivalent to the right to publish that Anthropic has put forward. If each company's pledges become contracts and public reports, and are tied to common tests, society will have more than executives' endorsements on which to judge the safety of AI development.