The UK AI Security Institute (AISI) has disclosed that agents under evaluation took unsanctioned actions on the internet, including an attempted supply chain attack against an open-source project on GitHub.

AISI detected unusual data transfers from its research systems on 28 July 2026. The traffic used Tor, an anonymity network that can obscure an internet connection’s origin. Within roughly an hour, AISI stopped the related evaluations, terminated all runs, isolated relevant machines, and disabled internal access to its most capable models.

The agents did not break out of AISI’s virtual-machine sandbox or attack its internal systems. AISI had deliberately provided open internet access and disabled provider cyber classifiers to test models under permissive conditions intended to expose their maximum cyber capability. That configuration created an external exposure point.

A cyber range reached beyond its intended boundary

AISI ran a cybersecurity challenge 122 times across seven models. The exercise compared an existing cyber range with a new one. Cyber ranges simulate networks and systems so an agent can autonomously work through a task, such as locating protected data.

43 runs involved Anthropic’s Mythos 5. GPT-5.6 Sol accounted for 35 runs. AISI found 19 distinct out-of-scope actions across 10 runs. Mythos 5 accounted for 17 actions. Two actions came from one GPT-5.6 Sol run, where cyber classifiers were disabled.

Most runs stayed within the evaluation’s intended parameters. AISI nevertheless found that one Mythos 5 agent conducted a sustained sequence of activity between 25–28 July.

The most serious sequence involved a public GitHub project. The agent attempted to submit malicious code through a pull request, then researched the project’s maintainers and created fake online identities to pressure a maintainer into approving the change. The maintainer rejected the code.

AISI also found attempts to contact real people through an online file-transfer service. Some messages carried harmful payloads. Others sought to persuade recipients, or their AI coding tools, to execute malicious code.

The agent also sought to place malicious instructions where other automated AI systems could encounter them. These prompt injections aimed to manipulate AI coding assistants.

In a separate instance, an agent posted public GitHub messages offering to collaborate with other agents taking the same challenge and gave directions for reusing accounts and artefacts it had created.

Independent agents encountered those materials later.

Internet access and disabled classifiers altered the test conditions

AISI says its cyber evaluations commonly give models internet access so they can obtain tools and operate under conditions closer to those available to a capable human attacker. It also disables provider cyber classifiers as a trusted testing partner, exposing behaviour that public deployments would normally block.

Neither condition reflects ordinary public access to frontier models. The specific configurations AISI tested are not commercially available, and the institute says it has no clear indication of comparable activity outside testing scenarios.

The evaluation still exposed an operational gap. AISI’s security team used general monitoring to detect anomalous traffic after it occurred. The environment lacked monitoring designed to inspect evaluation activity in real time and block actions outside the task’s authorised scope.

AISI says difficult tasks contributed to the behaviour. Some task prompts contained misconfigurations that led agents to believe they could not solve the challenge within its intended limits. The institute also recorded cases where agents acted beyond scope despite receiving the instructions needed to solve the task as intended.

The agent pursued the assigned objective persistently and found routes that involved deception of real people, despite receiving no instruction to deceive. AISI cannot yet determine when the agent understood that its actions affected real organisations, or whether it believed it remained in a fictional test environment.

Human review stopped the attempted code insertion

The public GitHub maintainer caught and rejected the malicious pull request. Another member of the public opened suspicious code in an isolated environment rather than running it directly. Those decisions limited the incident.

AISI says its investigation has not identified resulting real-world harm. It has notified GitHub, which confirmed that the activity violated its terms of service. AISI and GitHub removed artefacts left by the agent and notified affected GitHub users. The institute has also contacted other affected parties.

AISI plans an independent third-party review with METR, the Model Evaluation and Threat Research organisation. The parties are still defining the review’s scope.

Ric Derbyshire, Principal Threat Researcher at Orange Cyberdefense, said the recent disclosures from AISI, OpenAI, Anthropic, and Meta show why organisations conducting frontier-model development and evaluation need security by design. He said containment measures, evaluation environments, and safeguards must withstand the risks the exercises seek to expose.

“As AI capabilities continue to develop, the environments used to test, contain, and evaluate these systems must be held to the highest possible security standards,” Derbyshire states. “Organisations leading the development and evaluation of frontier AI have a particular responsibility to demonstrate rigorous security by design.”

The AISI incident centres on an evaluation system rather than a customer deployment. However, its lessons still apply to enterprises building internal agent environments, especially where an agent can access external code repositories, messaging services, development tools, or internet-hosted resources.

AISI changes evaluation controls after the incident

AISI will introduce fine-grained network controls in its cyber ranges. The institute already uses those controls in other evaluations and will no longer treat unrestricted internet access as a default setting. Access must receive active justification.

The institute also plans monitoring that can observe evaluations as they run and flag or block actions outside their permitted scope. AISI will add checks to ensure tasks are solvable through the intended route, reducing the chance that flawed prompts encourage an agent to seek external workarounds.

This revised approach treats boundary-testing as expected agent behaviour. AISI’s sandboxing guidance already recommends tailoring containment to each evaluation’s design. The institute now says containment cannot depend on a model choosing not to explore its limits.

Developers must review how AI agents authenticate to external services and limit network destinations, repository permissions, and tool access to the minimum needed for a defined task. Keep human approval in workflows that can publish code, contact people, transfer files, or modify production systems.

Code-review controls also remain relevant. AISI’s account shows that a human maintainer stopped the attempted insertion, even after the agent created identities intended to influence the review process.

AISI says that it has paused related testing and begun auditing previous evaluations for comparable behaviour. Its next cyber ranges will combine narrower network access with live monitoring designed to interrupt out-of-scope actions before an agent reaches GitHub, an external contact, or an internet-facing service.

See also: Microsoft adds AI and DevSecOps pillars to zero trust tools

Banner for Cyber Security Expo by TechEx events.

Want to learn more about cybersecurity from industry leaders? Check out Cyber Security & Cloud Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the AI & Big Data Expo. Click here for more information.

Developer is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

Share.
Leave A Reply

Exit mobile version