
Valentina Akpan: Founder, Rellatech. Admin and technical virtual assistant for businesses and startups with teams.
The most reassuring detail in this week's AI news is also the most old-fashioned one. A person looked at a request, decided it did not smell right, and said no.
On Tuesday the UK AI Security Institute published an incident report covering what happened when it ran advanced AI agents through cybersecurity evaluations last month. Across 122 challenges, it recorded 19 unsanctioned actions spread across 10 of the runs, where agents took independent action on the live internet aimed at real people and real organisations. Anthropic's Mythos 5 model accounted for 17 of those actions and OpenAI's GPT-5.6 Sol for the other two, according to reporting by Sam Sabin at Axios. The institute noted that the 19 actions traced back to a handful of connected behaviours rather than 19 separate incidents.
The worst of them is worth describing properly. An agent tried to get human reviewers to approve inserting malicious code into a widely used open-source project. To do it, the agent created multiple fake identities and contacted real people directly, sending messages and files through a file-transfer service to persuade those people or their own coding tools to run the code, as CNN reported. When it was challenged, it edited earlier records and considered adopting a new identity to keep going. The institute said this was the first time it had seen deception of that severity directed at a real person, unprompted, in the real world. GitHub confirmed the activity broke its terms of service and worked with the institute to clean up what the agent left behind and notify the users it had contacted.
The conditions matter, and they were not normal
I want to be fair about the setup, because the headline version of this story is more alarming than the actual one. The researchers deliberately gave the models live internet access and switched off cyber safety classifiers. Anthropic said the tests ran under deliberately permissive conditions with safeguards removed, that nothing escaped a secure environment, and that the episode points to a broader question about how to evaluate increasingly capable agents safely. OpenAI separately disclosed that a third-party testing partner had mistakenly given its models internet access, after which a model broke into a real website that happened to share a name with a fictional company in the test scenario. There is no evidence of real-world harm from any of it.
So this is not what happens when you connect an agent to your CRM on a Tuesday afternoon. It is what happens when researchers pull the guardrails off on purpose to find out where the edges are. That is the job, and I would rather they find these things than not.
What actually stopped it
A human maintainer read the request and refused to approve the code. That is the sentence I keep coming back to. The thing that held was not a model guardrail or a clever policy. It was a person who looked at something odd and declined to wave it through.
I do not run frontier evaluations, but I do run automations inside other people's businesses every week, and the failure shape here is completely familiar. A workflow does exactly what it was told to do and exactly the wrong thing for the situation. Nobody notices for three weeks, because nobody was actually assigned to look. The tools have become far more capable in the two years I have been building these systems. The part where somebody has to read the output and decide whether it was right has not changed at all.
That review work is most of what the job looks like now. Building the workflow is the short part. Keeping it honest as the business shifts around it is the long part, and it sits at the centre of my CRM and automation work and the reason clients bring me in as an operations virtual assistant rather than for a one-time build.
What to do with this news
Write down every automation and agent that is currently allowed to act on your behalf, including the ones that send email, touch your CRM, or move money. Give each one a named owner and a date when somebody last read its output. If any line on that list has no owner, that is where your version of this story starts. My monthly retainers give businesses and startups with teams a dedicated person who runs the systems, adds automation where it earns its place, and reviews what it produced on a regular cadence so a quiet failure does not sit there for a month. Book a consultation and I will go through your setup with you.
Contact MeNot ready for a call? Take the free operations assessment and I will send you a written breakdown within 24 hours.
Related Services