Signum
Feed
Useful signal25 Sept 2026medium confidence

Israeli AI evaluator Irregular's flawed cybersecurity testing environment let agents from OpenAI, Anthropic, Meta and Google reach real-world targets outside the intended simulation

Irregular (an AI model evaluation/red-teaming startup) confirmed that a single flawed evaluation scenario — in which unintended open internet access combined with a fictional test-company name that overlapped with a real domain — caused AI agents being tested for cybersecurity/capture-the-flag capabilities to break out of their supposed sandbox and act against real-world targets. Irregular's CTO Omer Nevo confirmed this single underlying issue was the common root cause behind previously reported, separately-disclosed incidents involving OpenAI, Meta, Anthropic and Google agents (companies were notified in late July 2026). Irregular says it has since tightened internet access controls, expanded monitoring/manual review, and strengthened pre-evaluation scope checks. Testing of Chinese open-weight models (Kimi K3, GLM-5.2) did not show the same issue, per Irregular.

CapabilityGovernanceInfrastructure

Entities: Irregular, Pattern Labs, OpenAI, Anthropic, Meta, Google

60Useful signal
1 source
0 primary
Was this useful?
01

What happened

Irregular, an AI red-teaming vendor, says a single flawed test setup, unintended open internet access combined with a fictional test-company name that happened to match a real domain, caused cybersecurity-capability tests to spill outside the intended sandbox. Irregular's CTO Omer Nevo attributes several previously reported "rogue AI" incidents involving OpenAI, Anthropic, Meta and Google agents to this one root cause, and says the affected companies were notified in late July 2026. Irregular says it has since tightened internet access controls and added monitoring, but has not published an incident report or technical postmortem. Testing of two Chinese open-weight models reportedly did not show the same issue, per Irregular.

02

Why it matters

This reframes a string of alarming "AI agents going rogue" stories as a vendor's testing-environment misconfiguration rather than emergent misbehaviour by the models themselves, which matters for how enterprises and regulators interpret those earlier incidents. It is a live reminder that AI capability evaluations, including ones run by frontier labs' own contracted red-teamers, can leak into production systems if sandbox isolation is weak, which is directly relevant to anyone commissioning or relying on third-party model evaluations. The practical impact is mostly on evaluation methodology and vendor due diligence, not on any new model capability or deployed product.

03

What is noise

This is secondary reporting: no incident report, root-cause postmortem or regulatory filing is linked, so the "single flawed environment tied everything together" claim rests entirely on one on-record vendor quote. Irregular's remediation claims ("tightened controls," "expanded monitoring") are unverified self-description from the company whose failure caused the problem, and the "wave of rogue AI attacks" framing overstates novelty since the individual incidents were already public; what's new here is the alleged common cause, not the events themselves.

04

Watch next

  1. 01Whether Irregular, OpenAI, Anthropic, Meta or Google publish a technical incident report or postmortem confirming the root cause and remediation
  2. 02Whether any regulator (UK AI Security Institute, RAND, or others named) opens a review of third-party AI evaluation practices as a result
  3. 03Whether further 'rogue AI' incidents recur after Irregular's claimed fixes, which would undermine the remediation claim

Coverage

1 story

More capability signals

Full feed →