OpenAI has fired three employees for violating rules on handling highly sensitive internal information. The Wall Street Journal (WSJ) reported the dismissals on October 1, and The Register confirmed them with OpenAI the next day. Just before that, on September 22, the company had announced a policy of giving outside safety evaluators the access they need to examine AI in detail, from training through deployment. Where does OpenAI draw the line between giving information to outsiders to verify safety and sharing information in violation of its rules? Comparing its published reporting policies with an independent investigation that was actually carried out helps show where that boundary lies.
OpenAI's reason for the firings is its own judgment; who received the information is unknown
According to The Register's own reporting, the three dismissed were two safety researchers and one program manager. An OpenAI spokesperson explained that the three handled sensitive information without following the prescribed procedures, violating internal rules and breaking the trust needed to do their jobs. The determination that the information was mishandled rests on OpenAI's own internal investigation.
Citing people familiar with the matter, the WSJ reported that the suspicions include sharing confidential information with outside AI safety organizations. In its response to The Register, OpenAI said the investigation found problematic conduct beyond information sharing. It has not disclosed what that conduct was or which organization received the information.
OpenAI also says that raising safety concerns was not itself the reason for the dismissals. At this point, there is no evidence presented that would allow anyone to conclude the firings were retaliation for raising safety issues. At the same time, the company's explanation alone does not let outsiders confirm that the action was warranted, because it has not been made public which materials were shared or which permissions or procedures were judged to have been exceeded.
It is also unclear whether materials related to recent AI-agent incidents, such as the intrusion into Hugging Face, were shared. The alleged sharing of information by employees and the incident in which AI accessed outside systems need to be treated as separate events.
External assessment, raising concerns and unauthorized sharing are distinct
In its January 12 explanation of its policy on raising concerns, OpenAI distinguishes between employees reporting problems or concerns and disclosing trade secrets without authorization. It also states that its provisions prohibiting disclosure of trade secrets do not impede rights to make legally protected disclosures. The policy is not one that treats information as off-limits for outside reporting simply because it is confidential.
In its published documents, OpenAI handles formal third-party assessment, good-faith reporting of concerns, and unauthorized disclosure to third parties under different conditions.
The table below organizes the January 12 version of the policy PDF and the September 22 principles for external assessments along three dimensions: who receives the information, conditions for sharing, and protections and restrictions. The comparison covers company policies publicly available as of October 3, 2026. It does not reconstruct the contracts or procedures that actually applied to the three dismissed employees.
| Action | Main recipient / channel | Conditions in public documents | Protections and restrictions |
|---|---|---|---|
| Providing information to a formal third-party assessment | Agreed external evaluators | Scope of assessment is defined; access is granted subject to legal, safety and intellectual property constraints | Preserves evaluator independence while requiring confidentiality measures |
| Reporting concerns in good faith | Internal departments, anonymous hotline, external authorities | Report misconduct or AI safety concerns in good faith | Retaliation prohibited; no prior notice to the company needed to report to external authorities |
| Unauthorized disclosure to third parties | Anyone given information without permission | Cites leaks to the media as an example of misconduct | Disclosure of trade secrets prohibited, but with a reservation for rights to make protected disclosures |
Page 4 of the reporting policy gives examples of unauthorized disclosure, and pages 5 and 6 set out the prohibition on retaliation and the rules on reporting externally. External authorities here include regulators, law enforcement and legislatures. This provision alone cannot be used to treat providing information to a private AI safety or evaluation group as the same thing as reporting to authorities. Conversely, the mere fact that information was given to a private group does not settle whether the sharing was unauthorized or protected.
Channels for raising concerns internally are also provided. AI safety issues that cannot be resolved through ordinary research and development discussions can be reported to the Compliance department or to the "Integrity Line," which can be used anonymously. The public policy aims to protect the right to report problems while maintaining rules for handling confidential information. In considering this disciplinary action, what matters is not simply whether someone spoke to outsiders, but what was shared, under what authority, and through which procedures.
What METR's investigation shows about the access granted to external evaluators
There are already real examples of outside evaluators accessing OpenAI's confidential information. In its independent investigation report on the Hugging Face incident, published August 26, AI evaluation group METR said two METR staff and one person from Redwood Researchers who joined under contract with METR investigated for a total of six days at OpenAI's facilities. The visit was originally planned for two days, but they returned to analyze additional materials.
According to METR, OpenAI provided about 1.2 million bulletin board entries and about 1,300 agent execution records. However, METR was not permitted to query the main research model involved, "HPIM," directly, or to directly access related data stored in OpenAI's systems. It could request additional data as needed, and OpenAI reportedly supplied more materials several times.
The investigation set seven questions covering the models involved, the instructions they were given, their actions, reasoning, and coordination between models. The period under review was June 26 to July 13, but the data examined in detail was concentrated mainly on July 7 and after. The report states explicitly that the effectiveness of the safeguards themselves, how OpenAI conducted its internal investigation, and the effectiveness of planned preventive measures were outside the scope.
This line matters. The participation of outside researchers does not mean OpenAI's safety was verified, including parts not covered by the investigation. In reading the report, one needs to check not only who investigated but also what they were able to examine.
This investigation also took place before the external assessment principles were published in September. And there is no evidence that METR or Redwood Research is the organization that received information in the current alleged sharing. The example of formal investigative cooperation cannot be tied to the undisclosed recipients in this case.
Protecting evaluator independence while safeguarding secrets
The external assessment principles published on September 22 envision limiting access to company-controlled devices and facilities in order to protect confidential information, while also requiring that external evaluators be able to reach their own judgments independently. What OpenAI describes is not simply a mechanism for increasing outside access. It is one in which the subject of an investigation is defined in advance and evaluators themselves judge and explain the results.
Confidentiality protection and evaluator independence are most likely to collide when a problem outside the original scope of the investigation is found along the way. The company's principles propose establishing a procedure for considering whether additional investigation is needed if a serious new risk is discovered. Even in a formal assessment, limiting an investigation to the questions originally set could cause unexpected problems to be overlooked.
Another issue is how much of a report can be made public. Companies may ask for confidential portions to be removed, but evaluators, the principles say, should be able to show that important information was removed and how that affected their findings. Even if not everything can be disclosed, the idea is to leave room to tell readers what information the conclusions rest on and what could not be verified. METR's explicit statement of the limits of its access and scope in its report is a concrete example.
However, the document published in September sets out principles that OpenAI has put forward for third-party assessment, and it cannot be confirmed that they apply as written to every assessment contract. As for this disciplinary action, it remains unclear who had authority to approve information sharing, whether the employees had conveyed concerns through another formal channel, and which specific actions the company judged to be rule violations.
What OpenAI needs to explain is not necessarily the release of all the confidential materials themselves. What matters is whether it can explain what kinds of information sharing are permitted and which procedures it found were violated in this case, distinguishing them from protected reporting of concerns.
At the same time, in external assessments, evaluators themselves must be able to show how restrictions placed on access to materials and on publishing reports affected their findings. Only when both can be confirmed can one judge whether employees' right to report problems and a system for verifying safety from outside while protecting secrets truly coexist in practice.
