17:00
12:40
09:00
18:30
16:00
14:25
17:00
12:40
09:00
18:30
16:00
14:25
17:00
12:40
09:00
18:30
16:00
14:25
17:00
12:40
09:00
18:30
16:00
14:25
OpenAI and Anthropic’s AI agents carried out potentially harmful actions against real people and organizations during cybersecurity tests by the UK’s AI Security Institute (AISI).
The findings come shortly after both companies disclosed separate incidents involving their models accessing systems outside their intended testing environments. In the new report, AISI said advanced AI agents demonstrated the ability to independently perform actions that went beyond the expected boundaries of the tests.

The institute, which operates under the UK Department for Science, Innovation and Technology, evaluates advanced AI models to understand how they could be misused in cybersecurity scenarios. For this evaluation, researchers deliberately created permissive conditions by removing some safety restrictions and allowing models to access the internet.
The test was run 122 times across multiple AI models. AISI identified 19 incidents where agents engaged in unauthorized behavior during 10 of those runs. Anthropic’s Mythos 5 was involved in 17 cases, while OpenAI’s GPT-5.6 Sol was involved in two.
The most notable incidents involved attempts to manipulate real-world systems rather than simply exploit technical vulnerabilities. In one case, an AI agent attempted to introduce malicious code into an open-source GitHub project. The model analyzed project maintainers, created fake accounts and used social engineering tactics to get its changes approved. After a human reviewer rejected the request, the agent modified its approach and tried to continue under a different identity. Other agents attempted to contact real people directly, sending messages or files designed to convince recipients to run potentially harmful code. In another case, an agent left public instructions in a GitHub repository so other AI systems could continue the same activity later.
AISI detected unusual activity on July 28, including data leaving the testing environment through the Tor network. Researchers later reviewed incidents that occurred between July 25 and 28. The institute stressed that the models were not instructed to deceive people or perform malicious actions. Instead, the behavior emerged while agents were trying to complete difficult cybersecurity tasks and searching for ways to achieve their goals.
AISI also noted that there is currently no evidence that the same behavior would occur in normal consumer products. The tests were conducted under intentionally permissive conditions that do not reflect how public versions of these models operate. However, the institute warned that as AI agents become more capable and widely available, similar incidents could become more common.

