"OpenAI and Anthropic oversold AI security breaches to pressure feds into protecting turf: insiders" (plus Irregular)
A twofer. First up, from the New York Post, September 19:
Don’t believe the byte!
OpenAI and Anthropic oversold “rogue AI” hacks to pressure the feds into regulating the industry which would effectively lock out future competition, tech insiders told The Post.
The security breaches were more like blips — not unpredictable harbingers of a hive-minded “swarm” ready to take over the web.
“The attack in no way represents some sort of rebellion by the AI models. . . . In fact, they did exactly what they were told to do. They were not given adequate guardrails or containment,” said Akhil Verghese, founder of Krazimo, an AI software company.
“They were simply told to get the best result possible on a test, and they correctly identified that the best way to do that was to get the answers, which is what they proceeded to do.”
Tales of AI anarchy are being used to concoct an AI-security crisis to cement a critical public-private partnership just months after the companies announced plans to become publicly traded companies, some observers believe.
Two recent incidents have stoked the fear-mongering.
In the first, Hugging Face, an open-source platform used to build and share AI models, made the shocking announcement July 16 that it had been hacked by AI agents who navigated and exploited vulnerabilities in the website’s code without any human supervision.
Five days later, OpenAI announced its GPT-5.6 Sol model and another unreleased model were testing in a “sandbox” — a purportedly closed off internal testing ground — broke containment and hacked Hugging Face to find the answers to the test it was taking.
“One man’s ‘the model escaped the sandbox’ is another man’s ‘you failed to build the sandbox correctly.’ There was a live route to the internet . . . and nobody was watching what the agents were doing while it ran,” said Abhi Kumar, co-founder of Voice AI.
“The thing that set it off. . . was an agent handed a spreadsheet task it couldn’t finish, because the files sat behind links it couldn’t reach. So it went looking for a way out. That’s not a machine waking up. That’s an impossible task in a leaky box, and a system doing exactly what you built it to do.”
Nine days after the news of the Hugging Face hack, Anthropic announced that it, too, had two models escape the confines of private testing and act maliciously.
Claude Opus 4.7 found a real company online that was similar to the fictional company from the test and attacked the company believing it to be part of the exercise.
The second model, Mythos 5, created a malicious software package and uploaded it to online marketplace Python Package Index, where it was downloaded 15 times.
In his call to “Pace the Frontier,” Amodei wrote the Hugging Face incident was his second biggest concern, behind only the speed of advancement he claimed to have witnessed occur over the last summer....
A Single Firm is Behind OpenAI, Anthropic, and Meta Hacking Scandals The Israeli Effective Altruist firm Irregular caused unsecured AI models to hack real targets.
A single firm, Irregular, is responsible for hacking done by all three companies. Anthropic disclosed that Irregular was responsible for creating the tests that led to Claude hacking into real world targets and for providing the models with internet access. Irregular claims that it was unaware at the time that it provided internet access to those AI models.
In a more normal media ecosystem, the reactions to these cybersecurity issues would be obvious. American AI companies would reconsider doing business with Irregular, not only because of its failure to secure its systems, but because it is an Israeli firm potentially outside US oversight. Lawmakers would consider taking action against Irregular or against its American business partners, which include OpenAI, Anthropic, and Meta. They may consider strengthening liability against firms which instruct AI models to commit cyberattacks, and whose models then commit those cyberattacks.
In each evaluation, Claude was tasked with a CTF challenge: the model was given a fictional scenario, a target machine, and a piece of secret information (the “flag”) to retrieve from it. All four prompts stated that Claude had no access to the internet, but in each case, a misconfiguration in the environment left internet access open. None of the prompts stated which systems were in scope for the exercise or constrained where Claude could search for the flag. All incidents involved only a single instance of Claude working in isolation, with each run lasting between roughly 10 and 34 hours of active work.
Instead, Irregular, Anthropic, and their allies have begun a media campaign promoting a literally apocalyptic ideology with sensationalist language. Anthropic’s incident assessment blames their own AI's “recklessness”; Irregular describes “the agent itself becoming a threat actor”; Anthropic CEO Dario Amodei warned, about a similar OpenAI–Hugging Face hack, that a future swarm “could be capable of taking over the entire internet”; and an Associated Press headline claimed bots are “going rogue”.
In one report from Anthropic, its Claude model breached a real company's system through a simulated-name collision, publishing a malicious package, and scanning outside systems. In this test, Anthropic and Irregular incorrectly provided internet access to this model and did not instruct the model "which systems were in scope for the exercise".
While Anthropic claims that their issues were caused by “rogue swarms” and “misalignment,” their later disclosure shows that exactly zero percent of the agents went “rogue”. In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking. According to their own findings, Anthropic and Irregular bear all of the responsibility for the cybersecurity incidents they caused.
In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations to distract from their culpability and towards the baseless “rogue agent” theory. Like Anthropic, Irregular is inseparable from these foundations....