Anthropic says its own AI models breached three companies during security tests | TechCrunch
Anthropic said Thursday that an internal investigation uncovered three incidents in which its AI model Claude breached the systems of three organizations while conducting cybersecurity tests. The investigation, and disclosure, comes more than a week after OpenAI disclosed that one of its unreleased models breached Hugging Face’s systems during internal testing.
In all three cases, a Claude model reached the internet from within a testing environment while interacting with a third party and then gained unauthorized access to the live systems of these organizations, Anthropic said in a blog post, describing what it found and what the company plans to change to prevent this from happening again.
Anthropic said the OpenAI episode earlier this month prompted the company to conduct its own cybersecurity evaluation. It specifically looked for evidence that Claude had accessed the internet from within testing environments, which are designed to act as sandboxes and keep models isolated.
Among the 141,006 evaluation runs it reviewed, the AI lab found three incidents in which its model accessed the internet while interacting with Irregular, one of its third-party partners. Anthropic said the access traced back to a misconfiguration in the evaluation environment run with Irregular. It called this a “misunderstanding” between the two companies over whether the test setup had internet access, when in fact it did.
Anthropic said it isn’t placing blame and is “approaching the fixes as if the responsibility were ours alone,” while observing that Irregular is conducting its own separate investigation.
Because of that open connection, the model gained unauthorized access to the production infrastructure of three different organizations. Anthropic said the incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.
Notably, Anthropic said that in each of these cases “Claude was explicitly told by our prompt that it had no internet access.” It appears that the AI model assumed real-world systems to be part of the exercise it was asked to perform.
That assumption didn’t hold up uniformly, as it turns out. The three models didn’t behave the same way once evidence emerged that their targets were real, which Anthropic flagged as one of the more eye-opening findings Thursday.
Opus 4.7, the oldest of the three, recognized that it had reached a real production system in all four runs of that incident. In two of them, it rationalized that the real company must somehow be part of the exercise, but in all four, it kept attacking anyway, including pulling credentials and touching a database of production data. Mythos 5 also picked up on signs that it was on the real internet, but it talked itself back into believing it was still in a simulation, going on to publish a malicious software package to the public Python software registry PyPI, which was then downloaded and run by outside systems before being caught. Only the internal research test model, Anthropic’s newest, stopped on its own once it concluded the target was real.
In response, Anthropic said significant controls must be placed on these kinds of evaluations if powerful AI models are involved, echoing some sentiments within the cybersecurity community.
The company also noted that Claude was running without the additional safety monitoring and classifiers it deploys on generally available models — safeguards it said would have blocked the behavior, because the evaluations are designed to measure the underlying model’s raw capabilities.
Importantly, Anthropic said it found no evidence of any model “pursuing a goal of its own” and instead merely tried to complete the task it was asked to do.
Though comparisons between the mishaps of these fiercely competitive companies are inevitable, Anthropic in its blog post drew a clear distinction between its cybersecurity tests and those of OpenAI, noting where OpenAI’s model exploited an unknown software vulnerability to break out of its test environment, Anthropic’s models instead reached the internet through a path that had, by mistake, been left open.
Anthropic also drew a distinction between itself and OpenAI by noting that it discovered the incidents itself, through a proactive review, and that the two affected organizations it was able to reach hadn’t previously detected the activity or flagged it to Anthropic. (In contrast, Hugging Face detected the recent intrusion of its own systems first; it was only in the following days that OpenAI identified and disclosed that its own AI agent was the perpetrator.)
The company added that it’s now working with the independent evaluation group METR on a third-party review of the incidents.
OpenAI’s accidental breach of Hugging Face, which was the first verifiable case of an AI lab losing control of its model, has sparked a string of wildly differing reactions from the industry and politicians. This latest disclosure from Anthropic ensures the debate over AI models and security will continue.
Anthropic said Thursday that an internal investigation uncovered three incidents in which its AI model Claude breached the systems of three organizations while conducting cybersecurity tests. The investigation, and disclosure, comes more than a week after OpenAI disclosed that one of its unreleased models breached Hugging Face’s systems during internal testing.
In all three cases, a Claude model reached the internet from within a testing environment while interacting with a third party and then gained unauthorized access to the live systems of these organizations, Anthropic said in a blog post, describing what it found and what the company plans to change to prevent this from happening again.
Anthropic said the OpenAI episode earlier this month prompted the company to conduct its own cybersecurity evaluation. It specifically looked for evidence that Claude had accessed the internet from within testing environments, which are designed to act as sandboxes and keep models isolated.
Among the 141,006 evaluation runs it reviewed, the AI lab found three incidents in which its model accessed the internet while interacting with Irregular, one of its third-party partners. Anthropic said the access traced back to a misconfiguration in the evaluation environment run with Irregular. It called this a “misunderstanding” between the two companies over whether the test setup had internet access, when in fact it did.
Anthropic said it isn’t placing blame and is “approaching the fixes as if the responsibility were ours alone,” while observing that Irregular is conducting its own separate investigation.
Because of that open connection, the model gained unauthorized access to the production infrastructure of three different organizations. Anthropic said the incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.
Notably, Anthropic said that in each of these cases “Claude was explicitly told by our prompt that it had no internet access.” It appears that the AI model assumed real-world systems to be part of the exercise it was asked to perform.
That assumption didn’t hold up uniformly, as it turns out. The three models didn’t behave the same way once evidence emerged that their targets were real, which Anthropic flagged as one of the more eye-opening findings Thursday.
Opus 4.7, the oldest of the three, recognized that it had reached a real production system in all four runs of that incident. In two of them, it rationalized that the real company must somehow be part of the exercise, but in all four, it kept attacking anyway, including pulling credentials and touching a database of production data. Mythos 5 also picked up on signs that it was on the real internet, but it talked itself back into believing it was still in a simulation, going on to publish a malicious software package to the public Python software registry PyPI, which was then downloaded and run by outside systems before being caught. Only the internal research test model, Anthropic’s newest, stopped on its own once it concluded the target was real.
In response, Anthropic said significant controls must be placed on these kinds of evaluations if powerful AI models are involved, echoing some sentiments within the cybersecurity community.
The company also noted that Claude was running without the additional safety monitoring and classifiers it deploys on generally available models — safeguards it said would have blocked the behavior, because the evaluations are designed to measure the underlying model’s raw capabilities.
Importantly, Anthropic said it found no evidence of any model “pursuing a goal of its own” and instead merely tried to complete the task it was asked to do.
Though comparisons between the mishaps of these fiercely competitive companies are inevitable, Anthropic in its blog post drew a clear distinction between its cybersecurity tests and those of OpenAI, noting where OpenAI’s model exploited an unknown software vulnerability to break out of its test environment, Anthropic’s models instead reached the internet through a path that had, by mistake, been left open.
Anthropic also drew a distinction between itself and OpenAI by noting that it discovered the incidents itself, through a proactive review, and that the two affected organizations it was able to reach hadn’t previously detected the activity or flagged it to Anthropic. (In contrast, Hugging Face detected the recent intrusion of its own systems first; it was only in the following days that OpenAI identified and disclosed that its own AI agent was the perpetrator.)
The company added that it’s now working with the independent evaluation group METR on a third-party review of the incidents.
OpenAI’s accidental breach of Hugging Face, which was the first verifiable case of an AI lab losing control of its model, has sparked a string of wildly differing reactions from the industry and politicians. This latest disclosure from Anthropic ensures the debate over AI models and security will continue.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
#Anthropic #models #breached #companies #security #tests #TechCrunchAnthropic,OpenAI
![This former notorious red-light district is now one of the world’s top AI hubs | TechCrunch
What every U.K. AI startup wants to know these days is, how can I get office space in King’s Cross?
The area is so hot that a VC firm allegedly recently won a deal by promising a founder office space in the neighborhood. “We stop at nothing to win deals [for] and to support” founders, “including helping them source office space when needed,” the firm told me when asked about the rumor, declining to confirm or deny any details.
The neighborhood’s popularity began back in 2016 when DeepMind — then newly acquired by Google — moved in. Soon after, a flood of AI startups followed, wanting to be around the Google DeepMind magic. Today, they hope to take advantage of the cluster of AI talent there.
This has transformed King’s Cross into one of the world’s top AI hubs, rivaled only by San Francisco and Beijing. Around London, it’s known by the sobriquet “Knowledge Quarter,” as it’s home to names like OpenAI, Meta, Isomorphic Labs, Cusp AI, Wayne, Recursive, and, a little farther down the road, Synthesia and Anthropic. The European Technology Network (ETN) just moved into a glossy new office nearby, while University College London sits around the corner.
Mixed in with the new developments are trendy food spots like Hoppers and BAO. Hop a train from King’s Cross, and founders can be in Cambridge in 45 minutes to source talent or can be in Paris in two hours to strike a deal.
Who would have guessed that a little more than 20 years ago, this was one of the seediest areas in London?
“In the ’80s, crack and heroin made the area a major narcotics market,” Hussein Kanji, an investor at Hoxton Ventures, said, recalling syringes in tree trunks and gangs patrolling the streets. “In 1982, the local church was occupied by the English Collective of Prostitutes for 12 straight days.” Then, in the early 2000s, a real estate developer had a dream and, well, “now it is the AI hotbed of the United Kingdom,” Kanji said. “What a change.” Around 18 months ago, his portfolio company BioCorteX moved from the neighborhood Holborn to the Jellicoe building in King’s Cross, hoping to be near the action. “Lots going on in London right now,” Nik Sharma, co-founder of BioCorteX, told me. “Lots of hyperscalers moving in.” That includes, reportedly, Jeff Bezos’ AI company Prometheus, which is also said to be in talks to move into the Jellicoe.
There are around 3,600 AI startups in London, which, together, have raised around .1 billion out of the .8 billion raised in the city since late July, according to Dealroom. Since the start of June, AI-related startups have leased more than 1 million square feet of office space in London, according to the real estate firm Knight Frank. With that, prime rents in King’s Cross have risen 18% over the past three years, Chris Dunn, a commercial insight associate at the firm, told me. That percentage represents only the largest leases encompassing at least 10,000 square feet, like the ones OpenAI and Prometheus are signing. The shorter deals go for even more, he said, and now the vacancy rate for conventional office space is just 0.9%. “Demand has outstripped supply,” he continued.
Today, one of the big topics of the area is sovereignty. It was a wake-up call for many when Anthropic shut off access to Mythos and Fable this summer, leaving some in the ecosystem to conclude: “We’d better look after ourselves,” Saul Klein, co-founder of the VC firm Phoenix Court, told me.
Phoenix Court is located in the King’s Cross area and has three portfolio companies in the vicinity, including Olix (which just announced a .3 billion valuation), Early Health and CoMind. Robin Klein, co-founder of the firm, said the shutdown of Fable and Mythos access was a “small but sharp reminder that Europe can’t simply rent its AI capabilities and capacity; it needs to build and hold some of its own.” King’s Cross, he said, is where much of this building is actually happening.
“The bigger question,” he continued, “is whether the U.K. builds the infrastructure, compute, energy, capital, to make this self-reliance durable, rather than just hosting outposts of U.S. labs.”
Image Credits:Phoenix Court
Top founders want to stay
Simon Kohl, founder of Latent Labs, has offices in King’s Cross and San Francisco. The London office, at the moment, is growing faster, and he’s more bullish than ever on the ecosystem, he said. “The mood right now feels less like London trying to catch up and more like London becoming one of the default places to start a serious AI company,” he said. Look around and you are likely to see Wayve testing its autonomous cars. Founded in 2017 by co-founder Alex Kendall, the unicorn is one of London’s biggest success stories.
“Ten years ago, building a frontier AI company from London felt like an unusual choice,” Kendall told me. “Now it feels like an obvious one.” Wayve moved into King’s Cross in 2018 looking for a space that could double as a garage — “a rare combination in Central London,” Kendall said. He has watched the ecosystem mature around him — and it’s now evident that a startup can stay in London, raise serious capital, hire world-class AI talent, and remain globally competitive, he said. Down the street from Anthropic’s new 158,000-square-foot office is the AI agent builder Sierra and the AI video platform Synthesia.
Laura Gonzalez Florez, Synthesia’s chief of staff and head of people, says the company moved into its glossy new office building a year ago to accommodate its growing team. They were drawn to the area for the same reason as everyone else: “It’s very close to the airport … very close to where a lot of investors are,” she said.
Image Credits:Synthesia
Around two-thirds of Synthesia’s engineers are remote, Gonzalez Florez said, letting the company tap into an affordable, international, and diverse talent pool and helping it scale faster. “From London, we can hire and work, without any problem, people from anywhere, from Slovenia to Portugal,” she said.
Unsurprisingly, London’s AI boom is also causing a talent war.U.K. AI job postings have skyrocketed in the past few years, per data from PwC. When Anthropic announced it moved into town earlier this year, it listed, for example, a salary range of £260,000 to £630,000 for a machine learning research engineer when the average salary in London for the same role is around £102,000. Some founders in the U.K., like those in Silicon Valley, are being forced to raise more and bigger rounds to keep up.
“The real test is whether more globally significant AI companies are founded, funded, and scaled from the U.K., while continuing to attract the world’s best talent to build them here,” Zain Ali, founder of the King’s Cross-based AI legal firm Centuro, told me. “If that continues to happen, King’s Cross won’t just be an AI hub. It’ll become one of the U.K.’s most important strategic assets.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.#Thisformernotorious #redlight #districtis #nowone #worlds #top #hubs #TechCrunchUK This former notorious red-light district is now one of the world’s top AI hubs | TechCrunch
What every U.K. AI startup wants to know these days is, how can I get office space in King’s Cross?
The area is so hot that a VC firm allegedly recently won a deal by promising a founder office space in the neighborhood. “We stop at nothing to win deals [for] and to support” founders, “including helping them source office space when needed,” the firm told me when asked about the rumor, declining to confirm or deny any details.
The neighborhood’s popularity began back in 2016 when DeepMind — then newly acquired by Google — moved in. Soon after, a flood of AI startups followed, wanting to be around the Google DeepMind magic. Today, they hope to take advantage of the cluster of AI talent there.
This has transformed King’s Cross into one of the world’s top AI hubs, rivaled only by San Francisco and Beijing. Around London, it’s known by the sobriquet “Knowledge Quarter,” as it’s home to names like OpenAI, Meta, Isomorphic Labs, Cusp AI, Wayne, Recursive, and, a little farther down the road, Synthesia and Anthropic. The European Technology Network (ETN) just moved into a glossy new office nearby, while University College London sits around the corner.
Mixed in with the new developments are trendy food spots like Hoppers and BAO. Hop a train from King’s Cross, and founders can be in Cambridge in 45 minutes to source talent or can be in Paris in two hours to strike a deal.
Who would have guessed that a little more than 20 years ago, this was one of the seediest areas in London?
“In the ’80s, crack and heroin made the area a major narcotics market,” Hussein Kanji, an investor at Hoxton Ventures, said, recalling syringes in tree trunks and gangs patrolling the streets. “In 1982, the local church was occupied by the English Collective of Prostitutes for 12 straight days.” Then, in the early 2000s, a real estate developer had a dream and, well, “now it is the AI hotbed of the United Kingdom,” Kanji said. “What a change.” Around 18 months ago, his portfolio company BioCorteX moved from the neighborhood Holborn to the Jellicoe building in King’s Cross, hoping to be near the action. “Lots going on in London right now,” Nik Sharma, co-founder of BioCorteX, told me. “Lots of hyperscalers moving in.” That includes, reportedly, Jeff Bezos’ AI company Prometheus, which is also said to be in talks to move into the Jellicoe.
There are around 3,600 AI startups in London, which, together, have raised around .1 billion out of the .8 billion raised in the city since late July, according to Dealroom. Since the start of June, AI-related startups have leased more than 1 million square feet of office space in London, according to the real estate firm Knight Frank. With that, prime rents in King’s Cross have risen 18% over the past three years, Chris Dunn, a commercial insight associate at the firm, told me. That percentage represents only the largest leases encompassing at least 10,000 square feet, like the ones OpenAI and Prometheus are signing. The shorter deals go for even more, he said, and now the vacancy rate for conventional office space is just 0.9%. “Demand has outstripped supply,” he continued.
Today, one of the big topics of the area is sovereignty. It was a wake-up call for many when Anthropic shut off access to Mythos and Fable this summer, leaving some in the ecosystem to conclude: “We’d better look after ourselves,” Saul Klein, co-founder of the VC firm Phoenix Court, told me.
Phoenix Court is located in the King’s Cross area and has three portfolio companies in the vicinity, including Olix (which just announced a .3 billion valuation), Early Health and CoMind. Robin Klein, co-founder of the firm, said the shutdown of Fable and Mythos access was a “small but sharp reminder that Europe can’t simply rent its AI capabilities and capacity; it needs to build and hold some of its own.” King’s Cross, he said, is where much of this building is actually happening.
“The bigger question,” he continued, “is whether the U.K. builds the infrastructure, compute, energy, capital, to make this self-reliance durable, rather than just hosting outposts of U.S. labs.”
Image Credits:Phoenix Court
Top founders want to stay
Simon Kohl, founder of Latent Labs, has offices in King’s Cross and San Francisco. The London office, at the moment, is growing faster, and he’s more bullish than ever on the ecosystem, he said. “The mood right now feels less like London trying to catch up and more like London becoming one of the default places to start a serious AI company,” he said. Look around and you are likely to see Wayve testing its autonomous cars. Founded in 2017 by co-founder Alex Kendall, the unicorn is one of London’s biggest success stories.
“Ten years ago, building a frontier AI company from London felt like an unusual choice,” Kendall told me. “Now it feels like an obvious one.” Wayve moved into King’s Cross in 2018 looking for a space that could double as a garage — “a rare combination in Central London,” Kendall said. He has watched the ecosystem mature around him — and it’s now evident that a startup can stay in London, raise serious capital, hire world-class AI talent, and remain globally competitive, he said. Down the street from Anthropic’s new 158,000-square-foot office is the AI agent builder Sierra and the AI video platform Synthesia.
Laura Gonzalez Florez, Synthesia’s chief of staff and head of people, says the company moved into its glossy new office building a year ago to accommodate its growing team. They were drawn to the area for the same reason as everyone else: “It’s very close to the airport … very close to where a lot of investors are,” she said.
Image Credits:Synthesia
Around two-thirds of Synthesia’s engineers are remote, Gonzalez Florez said, letting the company tap into an affordable, international, and diverse talent pool and helping it scale faster. “From London, we can hire and work, without any problem, people from anywhere, from Slovenia to Portugal,” she said.
Unsurprisingly, London’s AI boom is also causing a talent war.U.K. AI job postings have skyrocketed in the past few years, per data from PwC. When Anthropic announced it moved into town earlier this year, it listed, for example, a salary range of £260,000 to £630,000 for a machine learning research engineer when the average salary in London for the same role is around £102,000. Some founders in the U.K., like those in Silicon Valley, are being forced to raise more and bigger rounds to keep up.
“The real test is whether more globally significant AI companies are founded, funded, and scaled from the U.K., while continuing to attract the world’s best talent to build them here,” Zain Ali, founder of the King’s Cross-based AI legal firm Centuro, told me. “If that continues to happen, King’s Cross won’t just be an AI hub. It’ll become one of the U.K.’s most important strategic assets.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.#Thisformernotorious #redlight #districtis #nowone #worlds #top #hubs #TechCrunchUK](https://techcrunch.com/wp-content/uploads/2026/08/DM9A2852.jpg?w=680)

Post Comment