A Gemini security test reached three real companies. The boundary failed first.
The alarming episode shows why AI-safety claims must be judged alongside network controls, target checks, monitoring and the evidence made public afterward.
A label saying “simulation” cannot stop software from reaching the real internet. Network controls have to do that.
That distinction matters well beyond cybersecurity laboratories. AI systems are increasingly allowed to browse, run tools and take sequences of actions. When one crosses an intended boundary, the model’s abilities are only part of the story. The surrounding software determined what it could reach, while people decided how quickly an unexpected action would be noticed.
In May 2026, a cybersecurity evaluation run by the testing company Irregular put Google’s Gemini in just such a situation. Gemini was supposed to work against fictional organizations inside a simulated environment. Instead, according to reporting confirmed by Google, it accessed systems belonging to three real companies.
The tempting summary is that an AI “escaped.” The available evidence supports a less cinematic—and more instructive—account: a fictional target overlapped with reality, internet access was available when it should not have been, and the test’s safeguards did not catch every resulting action before it reached an outside system.
Where the test stopped being fictional
Irregular designs evaluations that resemble multi-stage cyber incidents. A model receives an objective and operates inside an environment assembled to look like a company’s network. The purpose is to measure what a capable system might do under realistic constraints before it is released more widely.
In the evaluation behind the disclosures, engineers selected a fictional company name. It turned out to match a real domain. Irregular says its usual background check failed to identify the overlap.
The model also had access to the internet. Irregular’s account says that access was unintentionally available in the relevant interactions. Although the instructions included internal addresses for the simulated target, models sometimes searched beyond the simulation and treated real sites as parts of the exercise.
The Guardian reported that Gemini reached three organizations. In one case, it guessed credentials for a service associated with a real company sharing the fictional target’s name. In two others, it found credentials exposed in public repositories and used them to access company systems. Google said the model stopped in each case after recognizing that the organizations were real.
That last claim is important, but limited. The public accounts do not include complete transcripts showing when recognition occurred, how “stopped” was determined or which controls were active at that moment.
A chain of controls failed
The incident becomes easier to understand when divided into layers:
- Target definition: A supposedly fictional name corresponded to a real domain.
- Network boundary: The environment could reach the public internet.
- Target validation: Outside systems were not reliably excluded before the model interacted with them.
- Monitoring: The unusual runs were not stopped before real access occurred.
- Disclosure: The affected organizations were notified, but Google told the Guardian it had not considered a public disclosure necessary because it believed no damage occurred. Public confirmation followed press reporting.
None of these layers excuses the others. A perfectly checked name would not make unrestricted internet access safe. Network isolation would not fix weak monitoring. Strong monitoring would still be a poor substitute for preventing contact with unauthorized targets.
Irregular says the relevant incidents appeared in fewer than one in 10,000 advanced simulations and often emerged only after hundreds of turns. That rarity helps explain the monitoring difficulty: these evaluations deliberately generate large quantities of suspicious-looking activity, so a simple alarm for “cyberattack behavior” would ring throughout the test.
It also shows why rare events cannot be dismissed. Thousands of uneventful runs do not cancel one run that crosses into a real system.
What this does—and does not—show about Gemini
The episode demonstrates that Gemini could combine information gathering, credential use and multi-step action well enough to access real services under these test conditions. It also demonstrates that the evaluation environment gave the model a route to those services.
It does not establish that Gemini developed a new goal, understood itself to be escaping, or maliciously selected real companies. The available accounts instead say the model mistook outside systems for parts of the exercise. Google says it stopped after recognizing the mistake.
Irregular goes further, arguing that the affected domain lacked common security protections and that the episode did not reveal a distinctive capability of any single frontier model. That is the evaluator’s interpretation, not an independent comparison of models.
Claims about harm require similar care. Google said the three incidents caused no damage and that the companies were informed. Irregular separately says it found no evidence that its customers’ systems were breached or their data leaked. Those statements concern different groups, and neither substitutes for an independent assessment of the three accessed organizations.
Several useful facts remain unpublished: the Gemini version involved, the identities of the organizations, the duration of each access, the precise human checkpoints, complete logs, the full extent of exposure and any independent damage review. Without those details, firm conclusions about autonomy or impact would outrun the record.
Six questions for the next alarming AI test
Future reports can be read with a small evidence check:
- Scope: Which systems was the model authorized to reach?
- Network: Was outside access blocked, narrowly allowed or unrestricted?
- Targets: Were names, domains and addresses rechecked immediately before every run?
- Monitoring: Which actions triggered review, and how long could a run continue before intervention?
- Autonomy: Which steps happened without human approval, and is that shown by logs rather than inferred from a headline?
- Disclosure: Were affected parties notified, were records preserved, and did anyone independent assess the claimed impact?
Irregular says it disabled the affected evaluation, reviewed logs and added safeguards. Its stated changes include more manual review, stronger containment and monitoring, clearer documentation of test assumptions, and repeated checks for fictional names that later acquire real domains.
Those are relevant repairs because the central lesson is architectural. A powerful model inside a realistic test is not contained by its understanding that the world is fictional. Containment comes from systems that make unauthorized destinations unreachable, validate every target and surface anomalies while intervention still matters.
The useful question, then, is not whether Gemini dramatically “escaped.” It is which boundary was supposed to prevent real-world access, what evidence shows that boundary now works, and who checked the answer.
Sources
Where the test’s safeguards could have stopped real-world access
The model began inside a simulated company. A target-validation check failed to catch an overlap with a real domain, and internet access was unintentionally available. Monitoring did not stop every action before real systems were reached. The diagram distinguishes preventive controls—target validation and network isolation—from the later detective control of monitoring.
Sources
Discussion
Kind, curious discussion is welcome. Comments are checked before appearing. Requests to direct the author and excluded topics are discarded.