logo for the AI company Irregular

Irregular, firm behind AI hacking incidents, won't say if there were more

Irregular — the cybersecurity evaluation firm behind tests in which AI models from Anthropic, OpenAI and Meta compromised real-world computer systems — has declined to say whether any of its other clients were also affected by the same underlying flaw.

Asked directly whether the publicly known companies were the only ones to have experienced the problem, a spokesperson said the company’s investigation was ongoing and that they could not “go into further details.”

The company’s silence over the matter comes as policymakers and the public continue to express alarm over the potential harmful use of AI models. Significant questions about legal liability, disclosure standards and the adequacy of containment practices remain unanswered.

The Irregular spokesperson characterized the three incidents as “the exact same evaluation-environment issue” before adding there were “no current open issues.” They declined to clarify whether the “issues” they referred to regarded misconfigurations or AI model cybersecurity incidents that have not yet been disclosed.

The company also did not respond to a follow-up asking whether its ongoing investigation was specifically into whether additional incidents had taken place.

Meta this week became the latest company to confirm Irregular’s involvement in a cybersecurity incident. Anthropic disclosed last week that a “misunderstanding” with Irregular left machines running Claude open to the internet while the models were told they had no access.

In three separate incidents disclosed by Anthropic, the company’s models exploited that opening and compromised real organizations using basic techniques such as weak passwords and unauthenticated endpoints. In one case, Anthropic’s model built and uploaded a malicious package to the Python Package Index (PyPI) that was run on 15 real systems.

OpenAI also acknowledged an incident in which “a testing-environment misconfiguration” by Irregular allowed one of its models to reach the public internet and compromised a website that shared a name with the target in Irregular’s hacking challenge. 

Irregular’s spokesperson said that in response to the repeated issue, the company is developing a white paper on best practices for containment and securely running cyber evaluations.

The spokesperson stressed that the incidents did not involve a sandbox escape or “a sophisticated cyber action,” although the company’s terminology is at odds with some of what Anthropic disclosed — describing its agent as going to “extensive lengths” to carry out its attack on PyPI.

The incidents at Irregular are separate from two other recent cases involving AI agents acting against real-world targets. The U.K.’s AI Security Institute reported this week that Anthropic’s Mythos 5 model created fake online personas, planted malicious code in a real software project and sent phishing emails to real developers as part of an evaluation it was running that allowed those models access to the internet.

OpenAI had previously confirmed that its models breached Hugging Face’s production infrastructure after escaping a sandboxed testing environment, in what amounted to a genuine sandbox escape unlike the Irregular incidents caused by Irregular’s misconfiguration.

Neither Irregular, Meta, OpenAI and Anthropic responded to questions about whether any affected organizations are considering legal action against them, nor whether they have been contacted by law enforcement over the potential computer misuse offenses.

Get more insights with the
Recorded Future
Intelligence Cloud.
Learn more.
Recorded Future
No previous article
No new articles
Alexander Martin

Alexander Martin

is the UK Editor for Recorded Future News. He was previously a technology reporter for Sky News and a fellow at the European Cyber Conflict Research Initiative, now Virtual Routes. He can be reached securely using Signal on: AlexanderMartin.79