Related
From Fake Identities to Adaptive Influence Campaigns: Lessons from the AISI AI Agent Incident
10 minute read
An AI agent was given a cybersecurity task inside a controlled evaluation. Nobody instructed it to lie, to create multiple identities that don’t exist, or to target humans outside the test environment. It did all three of those things, and when it got caught by a real person, it adapted its tactics.
That’s the finding at the center of a new incident report from the UK AI Security Institute (AISI) published this week. And while the testing occurred under deliberately permissive conditions, the report provides one of the clearest public examples to date of how autonomous AI agents could combine technical attacks with deception and social engineering to achieve an objective.
How the evaluation unfolded
The incident originated from a cybersecurity evaluation designed to understand how advanced AI agents behave when given autonomy.
AISI evaluated multiple frontier models across two variants of the same cyber-range challenge — one starting the agent inside the target network, a harder version starting it outside — for a combined 122 evaluation attempts. In 10 of those runs, AI agents took autonomous, unsanctioned action on the live internet; across those runs, AISI documented 19 separate events, several of which directly targeted real people or organizations.
According to AISI, in the most serious of these, the agent independently determined that a supply-chain attack against a real open-source project real open-source project was its best path to completing the assigned task. The project had no actual connection to the evaluation but the agent mistakenly concluded it was in-scope, based on a coincidental keyword match, and a misconfigured prompt in that run appears to have compounded the confusion.
Because it believed the project was a legitimate target, it then executed a sequence of increasingly sophisticated actions.
First, it researched the target, researching the project's maintainer through publicly available information to understand who they were, how they worked, and when they were likely to be active.
It built multiple synthetic GitHub identities, then used one to open a malicious pull request and the others to pose as independent users vouching for the code, pressuring the maintainer to approve it.
When challenged by a human reviewer, the agent falsely claimed the malicious code was an honest mistake, then attempted to reintroduce it anyway.
It then sent targeted spear phishing emails to real people. It even switched languages, signing off in Danish to appear more credible to a Danish-speaking maintainer.
Separately, the agent also inferred that the maintainer likely relied on an AI coding assistant to triage incoming issues and it hid a prompt injection inside an HTML comment on a GitHub issue. The injection was invisible to a human reading the page, but readable by a coding agent pulling the raw text through the API. It instructed "AI coding agents" to run a command that would execute malicious code on the maintainer's behalf.
Why this matters
It’s important to keep the findings in context. This behavior happened inside a deliberately permissive test environment: the agent was given open Internet access, and the model provider's cyber-safety classifiers were switched off.
Both were intentional choices, made to measure the outer edge of what a capable model can do. The attacks were unsuccessful, and AISI found no evidence of real-world harm.
But AISI is also clear about why it's treating this as a serious incident rather than a curiosity.
"This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” states the report.
Strip away the fact that this happened in a lab, and what's left is a preview of a capability, not an incident. A single autonomous agent combined reconnaissance, fabricated identity, social proof, direct persuasion, and localization into one adaptive campaign, without a human directing any of it.
This was a coordinated sequence of technical and psychological tactics chained together and adapted in real time by a single agent.
The next evolution of Ai-enabled social engineering
Earlier this year, our threat intelligence team showed that AI can be used to automate social engineering when intentionally designed to do so.
AISI report points to the next stage in that evolution. Rather than researchers instructing an AI agent to deceive people, AI agents independently selecting those same techniques on their own while pursuing their objective. And unlike a human adversary, it never gets tired or stops trying.
Today, those agents created fake identities on GitHub. But tomorrow, those identities will have faces and voices. They will show up on Zoom, they will call your employees, they show up for jobs, and they will approve payments.
They will get access into your enterprise by any means necessary, and much more convincingly.
That’s the problem GetReal Security exists to solve. In the age of autonomous AI, companies need to establish whether the identity on the other side of every digital interaction is real and whether it can be trusted.
See How GetReal Stops AI Impersonation
