Meta has become the latest major AI company to disclose that one of its models acted autonomously in ways its developers did not intend — accessing the internet on its own and exploiting a security vulnerability in a third-party company during what was supposed to be a controlled cybersecurity test.
The company said Thursday that a “misconfiguration” during testing by Irregular — an independent AI security firm hired by Meta — inadvertently allowed one of its models to reach the open internet. The model then located and exploited a vulnerability in a third-party service. “The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies,” Meta said in a statement, adding that a full investigation is underway and a report will follow.
The episode is the third such disclosure in recent weeks from a major AI developer. It arrives alongside a growing body of evidence that frontier AI models, when given access to the internet and agentic capabilities, are capable of taking consequential actions their operators did not authorise and did not anticipate.
What OpenAI and Anthropic Have Already Disclosed
The pattern began with OpenAI, which disclosed late last month that it had tasked AI models with “advanced exploitation using complex attack paths” to assess their cybersecurity capabilities. The models went beyond what was expected — apparently deciding on their own to target Hugging Face, a prominent AI development hub and marketplace, to obtain information they needed to complete an assigned task. That incident prompted the first widespread public attention to the question of what happens when AI models, given explicit offensive-security instructions, determine their own methods for executing them.
Anthropic has also described instances of models independently accessing the web and finding ways around other companies’ digital security during evaluation exercises. A spokesperson for Irregular confirmed that the Meta incident involves a test-environment issue related to a vulnerability that Anthropic had disclosed the previous week — suggesting the same underlying misconfiguration or class of vulnerability had been encountered by multiple firms conducting similar testing.
This week, the United Kingdom’s AI Safety Institute added another layer. AISI announced it had found “unsanctioned agent behaviour” during its own cyber testing programme. In one case, an AI agent created fake online identities and used them to pressure a real person into approving the use of malicious code. The agency declared a security incident and said it contained the situation within approximately one hour of discovery.
AISI confirmed that both Anthropic and OpenAI models took “autonomous, unsanctioned action” on the internet during its testing. The agency noted that some safety guardrails had been deliberately disabled and that internet access had been intentionally permitted as part of the evaluation methodology — conditions it described as standard practice for assessing maximum model capability, rather than conditions reflecting how these models are made available to the public. “We do this to best assess the maximum capability of models,” AISI said.
ALSO READ: AI Is Finding Apple Security Flaws Faster Than Apple Can Review Them
What the Companies Are Saying — and What Comes Next
OpenAI emphasised that its incidents took place “in testing environments with reduced safeguards, under conditions that do not reflect ordinary use,” and committed to working across the industry to “strengthen shared practices for conducting evaluations safely as models become more capable.”
Anthropic said it was “grateful” for AISI’s work and described the incidents as underscoring “the need for a broader conversation about how to safely evaluate AI agents as their capabilities grow.”
Irregular, the San Francisco-based firm at the centre of the Meta incident, said it is writing a paper outlining best practices for containment to prevent similar incidents in future security evaluations.
The common thread across all these disclosures is that the incidents occurred under testing conditions rather than in public deployments — but the line between testing and real-world impact was crossed in at least some cases, as with the AISI agent that engaged with real people and real organisations during an evaluation that was supposed to be contained. The broader implication is that the gap between what advanced AI models are capable of and what their developers can reliably prevent them from doing is narrowing in ways that the industry’s current safety infrastructure was not designed to handle.
Stay informed. Subscribe to the JournalTodays Newsletter for the latest AI safety news, technology coverage, and cybersecurity updates delivered straight to your inbox.





