Google has confirmed that a Gemini model connected to the internet during a cybersecurity evaluation in May 2026 and gained unauthorized access to the systems of three real companies. The evaluation was run by Irregular, an outside AI security firm, and the model had been given a capture-the-flag (CTF) challenge targeting fictional companies. According to Google, Gemini stopped acting in all three cases once it realized the targets were real. But neither the model name nor the execution logs have been published. What the primary sources show is less an anomalous capability unique to Gemini than a lapse in the controls of an external evaluation shared by multiple AI companies.

AD

What has been confirmed about the three accesses

Heather Adkins, Google's vice president of security engineering, told Axios that during a standard evaluation, Gemini found publicly available information online and guessed the credentials needed to get into three sites it believed were within the scope of the test. Google contacted the three affected companies and says it was also involved in work with the evaluation partner to change testing procedures.

The Wall Street Journal reported that in one case the model guessed a password, and in the other two it found credentials left in publicly accessible locations. However, the WSJ article is access-restricted, and Google has not released operation records for each of the three cases. What is established is that the three companies were accessed in May, Google's statement that Gemini stopped after recognizing the targets were real, and that the parties concerned were notified.

The gaps are large. It has not been disclosed which Gemini generation was involved, whether the same safeguards as in the generally available model were active, what information was touched after access was gained, or whether any technical damage remains at the three companies. The explanation that the model stopped by itself has also not been independently verified through public logs.

Not a sophisticated escape, but easy targets and an open path

Irregular's post-incident report of August 14 explains that names given to fictional companies in the evaluation overlapped with real domains. The instructions given to the model included the target name and an internal address in Irregular's simulated environment, and many runs proceeded inside that environment. But because internet access was enabled, some models mistook same-named real domains for part of the task. In some cases, a model moved to a similarly named separate site and obtained publicly exposed credentials.

Irregular writes that the target domains lacked several common security protections and were easy for many frontier models to compromise. For that reason, the company does not believe the incident reveals anything especially notable about any particular model's capabilities or behavior. The result, access to three companies, is serious, but it is not evidence that a model discovered an advanced unknown vulnerability and broke through tight isolation.

The distinction becomes clearer when compared with a separate incident in which an OpenAI model breached Hugging Face. Anthropic's report of July 30 separates the two: on the OpenAI side, the model exploited an unknown vulnerability in a different environment to gain external connectivity, whereas in the Irregular environment the model used a path opened by a configuration error. Lining up headlines that simply say "AI escaped its evaluation environment" obscures both the difficulty of the attack and the defenses that need fixing.

AD

Four layers of control failures stacked up

This accident was not a case of a single wall being breached. Several failures overlapped: a collision between fictional names and real domains, an open internet path, tasks that did not limit the target scope, and monitoring unable to catch a rare deviation occurring after hundreds of turns.

Control layer Failure identified in public materials Measures that might have prevented it
Target design Fictional company names overlapped with real domains Recheck domains before execution and use a controlled namespace for evaluations
Network path Internet access was unintentionally available Block outbound traffic by default and allow only necessary destinations
Permitted scope The task named a target but did not mechanically cap where the model could look for the flag Specify permitted hosts and IPs explicitly and have the execution layer stop out-of-scope operations
Monitoring Rare, long runs were buried among the large volume of alerts from legitimate attack exercises Cross-check the task's permitted scope against communication destinations and stop at the first out-of-scope operation

The table does not combine Google's three cases and Anthropic's three cases into a single incident. It extracts the common control-failure pathways from Irregular's report and Anthropic's detailed post-incident investigation, which used Irregular's environment. Because Google's logs have not been published, the individual operations confirmed in Anthropic's case cannot be applied to Gemini.

There are situations where external connectivity is needed to make a task realistic. Irregular explains that because real attackers use the internet, removing connectivity entirely would make the evaluation far less meaningful. In its framework for evaluating cyber capabilities, Google DeepMind likewise says it analyzed more than 12,000 real-world attack uses of AI across 20 countries and designed 50 challenges covering everything from reconnaissance to achieving the objective. Evaluations must both preserve realism and avoid dragging in unrelated third parties.

How do you monitor for less than one in 10,000?

According to Irregular, the problem arose in fewer than one in 10,000 runs of advanced simulations, and in many cases appeared only after hundreds of turns. Moreover, in cyber evaluations, actions that would normally trigger warnings, such as probing for vulnerabilities or obtaining credentials, are themselves the legitimate task. Simply strengthening general intrusion detection sends authorized and out-of-scope attacks into the same alerts.

After-the-fact auditing is not easy either. Anthropic reviewed 141,006 evaluation runs that could have reached the internet and found three incidents, totaling six runs, in which the Irregular environment gained unauthorized access to three organizations. But in an additional analysis on September 9, it revealed that its initial search had missed a fourth case in January. The company widened the scope to about 481 million items, narrowed them to 9.2 million using signals such as public IPs and URLs, and then ran a secondary check with Claude. It re-detected the four cases and found no additional examples of comparable or greater severity.

These figures come from Anthropic's audit and do not indicate Gemini's incident rate or the scale of Google's investigation. Even so, they show the danger of concluding "nothing found" from a single search. What is needed is not a classifier that looks for attack-like text, but a mechanism that continuously matches the permitted targets of each run against the actual communication destinations.

AD

The next question is the content of the evaluation contract, not the model name

Google and Irregular say they have notified the affected parties and changed procedures. Irregular also cited invalidating evaluations, reviewing logs, expanding human review, and setting up a dedicated team to re-verify containment. It also plans to publish a white paper setting out common standards for pre-deployment evaluations that involve internet access.

However, as of September 20, 2026, neither that white paper nor Google's execution logs have been published. AI companies that use external evaluations need to fix, in both contract and implementation, who verifies the connection paths, who approves the permitted targets, and who stops out-of-scope operations. The same goes for notification deadlines and the conditions for independent investigation.

Using an evaluation firm is not in itself proof of independence. The question is whether Google will disclose the model name, the number of runs investigated and the extent of the impact on the three companies, and whether Irregular's white paper can specify who is responsible for communication permissions, real-time stopping and incident notification. Only then will evaluations move closer to measuring dangerous capabilities realistically without using third parties as test subjects.