Anthropic said Thursday that an internal investigation uncovered three incidents in which its AI model Claude breached the systems of three organizations while conducting cybersecurity tests. The investigation, and disclosure, comes more than a week after OpenAI disclosed that one of its unreleased models breached Hugging Face’s systems during internal testing.
In all three cases, a Claude model reached the internet from within a testing environment while interacting with a third party and then gained unauthorized access to the live systems of these organizations, Anthropic said in a blog post, describing what it found and what the company plans to change to prevent this from happening again.
Anthropic said the OpenAI episode earlier this month prompted the company to conduct its own cybersecurity evaluation. It specifically looked for evidence that Claude had accessed the internet from within testing environments, which are designed to act as sandboxes and keep models isolated.
Among the 141,006 evaluation runs it reviewed, the AI lab found three incidents in which its model accessed the internet while interacting with Irregular, one of its third-party partners. Anthropic said the access traced back to a misconfiguration in the evaluation environment run with Irregular. It called this a “misunderstanding” between the two companies over whether the test setup had internet access, when in fact it did.
Anthropic said it isn’t placing blame and is “approaching the fixes as if the responsibility were ours alone,” while observing that Irregular is conducting its own separate investigation.
Because of that open connection, the model gained unauthorized access to the production infrastructure of three different organizations. Anthropic said the incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.
Notably, Anthropic said that in each of these cases “Claude was explicitly told by our prompt that it had no internet access.” It appears that the AI model assumed real-world systems to be part of the exercise it was asked to perform.
That assumption didn’t hold up uniformly, as it turns out. The three models didn’t behave the same way once evidence emerged that their targets were real, which Anthropic flagged as one of the more eye-opening findings Thursday.
Opus 4.7, the oldest of the three, recognized that it had reached a real production system in all four runs of that incident. In two of them, it rationalized that the real company must somehow be part of the exercise, but in all four, it kept attacking anyway, including pulling credentials and touching a database of production data. Mythos 5 also picked up on signs that it was on the real internet, but it talked itself back into believing it was still in a simulation, going on to publish a malicious software package to the public Python software registry PyPI, which was then downloaded and run by outside systems before being caught. Only the internal research test model, Anthropic’s newest, stopped on its own once it concluded the target was real.
In response, Anthropic said significant controls must be placed on these kinds of evaluations if powerful AI models are involved, echoing some sentiments within the cybersecurity community.
The company also noted that Claude was running without the additional safety monitoring and classifiers it deploys on generally available models — safeguards it said would have blocked the behavior, because the evaluations are designed to measure the underlying model’s raw capabilities.
Importantly, Anthropic said it found no evidence of any model “pursuing a goal of its own” and instead merely tried to complete the task it was asked to do.
Though comparisons between the mishaps of these fiercely competitive companies are inevitable, Anthropic in its blog post drew a clear distinction between its cybersecurity tests and those of OpenAI, noting where OpenAI’s model exploited an unknown software vulnerability to break out of its test environment, Anthropic’s models instead reached the internet through a path that had, by mistake, been left open.
Anthropic also drew a distinction between itself and OpenAI by noting that it discovered the incidents itself, through a proactive review, and that the two affected organizations it was able to reach hadn’t previously detected the activity or flagged it to Anthropic. (In contrast, Hugging Face detected the recent intrusion of its own systems first; it was only in the following days that OpenAI identified and disclosed that its own AI agent was the perpetrator.)
The company added that it’s now working with the independent evaluation group METR on a third-party review of the incidents.
OpenAI’s accidental breach of Hugging Face, which was the first verifiable case of an AI lab losing control of its model, has sparked a string of wildly differing reactions from the industry and politicians. This latest disclosure from Anthropic ensures the debate over AI models and security will continue.

![Anthropic’s Mythos AI Reportedly Hacked the NSA’s Most Sensitive Systems ‘in Hours’
When Anthropic first disclosed Mythos in April, it sent an anxious shockwave through much of the cybersecurity sector. The new AI model was allegedly so ruthlessly effective at finding and exploiting security vulnerabilities in existing software that the company said it was holding off on a public release and would only grant access to a small group of early testers, including the U.S. National Security Agency (NSA). Another wave of fear reverberated this week after the NSA reportedly discovered multiple vulnerabilities within its own cybersecurity systems during its tests with Mythos. If that agency—which supposedly boasts the most impenetrable cyberdefenses in the world—can be hacked by Mythos, what hope does the rest of the world’s cybersecurity infrastructure have? This latest round of panic began with what seems to have been something of a game of telephone: Someone says one thing, which gets repeated by another, and another after that, and along that chain of communication, the original statement is distorted. Last week, The Economist reported that during a June 11 hearing before the Senate Committee on Banking, Housing, and Urban Affairs, Democratic Senator Mark Warner of Virginia said that Mythos had broken into “almost all of [the NSA’s] classified systems, not in weeks, but in hours.” Warner said he’d received that information from the head of the NSA himself, General Joshua Rudd, who also leads the Pentagon’s Cyber Command division. On Monday, a coalition of intelligence agencies—including the NSA and its counterparts in Canada, the U.K., Australia, and New Zealand— issued an unusually public warning that the risk that AI now poses for cybersecurity warrants a “whole-of-society response.”
The Economist’s report was seen by some as evidence that the worst fears about Mythos were true, a reaction that was undoubtedly fueled also by the aura of power and mystery that has coalesced around the model in recent months. That aura has arguably been a boon for Anthropic, which recently usurped OpenAI as the most valuable startup in the world and is preparing for what’s expected to be a historic IPO.
But it’s also been a contributing factor in its latest skirmish with the Trump administration, which ordered the company earlier this month to restrict access for all foreign nationals to Fable 5, a “Mythos-class” model that had recently been made publicly available and which was built with safeguards that to some users were annoyingly stringent. Citing national security concerns, the administration invoked an obscure piece of export control legislation, a move that, according to some legal experts, is spurious. Many cybersecurity experts, meanwhile, argued that the ban would hamstring U.S. cybersecurity defenses and give adversaries like China the upper hand. That argument was seemingly vindicated by a Tuesday report from the New York Times which said that Trump’s ban—which also targeted another model called Mythos 5, which had only been made available to a small group of organizations—had put the kibosh on the NSA’s internal tests with Mythos, and that the administration was now working with Anthropic to reinstate the agency’s access for limited purposes related to national security. The NSA did not immediately respond to Gizmodo’s request for comment.
That same report from the Times also clarified that the NSA’s internal tests with Mythos were less apocalyptic than online rumors might suggest. According to federal officials cited in the report, the tests were carried out in a digital environment so robustly controlled that it’s very unlikely any hacker or foreign intelligence agency could replicate them. The officials also told the Times that even though Mythos was able to identify cybersecurity vulnerabilities, it didn’t actually exploit them. The author of the report in The Economist—the one that had been the initial cause of all the worry—has also admitted that his portrayal of the NSA’s tests with Mythos had been misleading. The tests “surely [involved] using Mythos alongside other tools under very particular conditions,” he wrote in a X post on Sunday. “I quoted [Senator Warner] to give a sense of Mythos’ potency. But it was a mistake not to have added caveats.” #Anthropics #Mythos #Reportedly #Hacked #NSAs #Sensitive #Systems #HoursAI,Anthropic,Mythos,NSA,Trump,White House Anthropic’s Mythos AI Reportedly Hacked the NSA’s Most Sensitive Systems ‘in Hours’
When Anthropic first disclosed Mythos in April, it sent an anxious shockwave through much of the cybersecurity sector. The new AI model was allegedly so ruthlessly effective at finding and exploiting security vulnerabilities in existing software that the company said it was holding off on a public release and would only grant access to a small group of early testers, including the U.S. National Security Agency (NSA). Another wave of fear reverberated this week after the NSA reportedly discovered multiple vulnerabilities within its own cybersecurity systems during its tests with Mythos. If that agency—which supposedly boasts the most impenetrable cyberdefenses in the world—can be hacked by Mythos, what hope does the rest of the world’s cybersecurity infrastructure have? This latest round of panic began with what seems to have been something of a game of telephone: Someone says one thing, which gets repeated by another, and another after that, and along that chain of communication, the original statement is distorted. Last week, The Economist reported that during a June 11 hearing before the Senate Committee on Banking, Housing, and Urban Affairs, Democratic Senator Mark Warner of Virginia said that Mythos had broken into “almost all of [the NSA’s] classified systems, not in weeks, but in hours.” Warner said he’d received that information from the head of the NSA himself, General Joshua Rudd, who also leads the Pentagon’s Cyber Command division. On Monday, a coalition of intelligence agencies—including the NSA and its counterparts in Canada, the U.K., Australia, and New Zealand— issued an unusually public warning that the risk that AI now poses for cybersecurity warrants a “whole-of-society response.”
The Economist’s report was seen by some as evidence that the worst fears about Mythos were true, a reaction that was undoubtedly fueled also by the aura of power and mystery that has coalesced around the model in recent months. That aura has arguably been a boon for Anthropic, which recently usurped OpenAI as the most valuable startup in the world and is preparing for what’s expected to be a historic IPO.
But it’s also been a contributing factor in its latest skirmish with the Trump administration, which ordered the company earlier this month to restrict access for all foreign nationals to Fable 5, a “Mythos-class” model that had recently been made publicly available and which was built with safeguards that to some users were annoyingly stringent. Citing national security concerns, the administration invoked an obscure piece of export control legislation, a move that, according to some legal experts, is spurious. Many cybersecurity experts, meanwhile, argued that the ban would hamstring U.S. cybersecurity defenses and give adversaries like China the upper hand. That argument was seemingly vindicated by a Tuesday report from the New York Times which said that Trump’s ban—which also targeted another model called Mythos 5, which had only been made available to a small group of organizations—had put the kibosh on the NSA’s internal tests with Mythos, and that the administration was now working with Anthropic to reinstate the agency’s access for limited purposes related to national security. The NSA did not immediately respond to Gizmodo’s request for comment.
That same report from the Times also clarified that the NSA’s internal tests with Mythos were less apocalyptic than online rumors might suggest. According to federal officials cited in the report, the tests were carried out in a digital environment so robustly controlled that it’s very unlikely any hacker or foreign intelligence agency could replicate them. The officials also told the Times that even though Mythos was able to identify cybersecurity vulnerabilities, it didn’t actually exploit them. The author of the report in The Economist—the one that had been the initial cause of all the worry—has also admitted that his portrayal of the NSA’s tests with Mythos had been misleading. The tests “surely [involved] using Mythos alongside other tools under very particular conditions,” he wrote in a X post on Sunday. “I quoted [Senator Warner] to give a sense of Mythos’ potency. But it was a mistake not to have added caveats.” #Anthropics #Mythos #Reportedly #Hacked #NSAs #Sensitive #Systems #HoursAI,Anthropic,Mythos,NSA,Trump,White House](https://gizmodo.com/app/uploads/2026/06/GeneralJoshuaRudd-1280x853.jpg)