×
Anthropic says its own AI models breached three companies during security tests | TechCrunch
Anthropic said Thursday that an internal investigation uncovered three incidents in which its AI model Claude breached the systems of three organizations while conducting cybersecurity tests. The investigation, and disclosure, comes more than a week after OpenAI disclosed that one of its unreleased models breached Hugging Face’s systems during internal testing.

In all three cases, a Claude model reached the internet from within a testing environment while interacting with a third party and then gained unauthorized access to the live systems of these organizations, Anthropic said in a blog post, describing what it found and what the company plans to change to prevent this from happening again.







Anthropic said the OpenAI episode earlier this month prompted the company to conduct its own cybersecurity evaluation. It specifically looked for evidence that Claude had accessed the internet from within testing environments, which are designed to act as sandboxes and keep models isolated.

Among the 141,006 evaluation runs it reviewed, the AI lab found three incidents in which its model accessed the internet while interacting with Irregular, one of its third-party partners. Anthropic said the access traced back to a misconfiguration in the evaluation environment run with Irregular. It called this a “misunderstanding” between the two companies over whether the test setup had internet access, when in fact it did. 

Anthropic said it isn’t placing blame and is “approaching the fixes as if the responsibility were ours alone,” while observing that Irregular is conducting its own separate investigation.

Because of that open connection, the model gained unauthorized access to the production infrastructure of three different organizations. Anthropic said the incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.

Notably, Anthropic said that in each of these cases “Claude was explicitly told by our prompt that it had no internet access.” It appears that the AI model assumed real-world systems to be part of the exercise it was asked to perform.


That assumption didn’t hold up uniformly, as it turns out. The three models didn’t behave the same way once evidence emerged that their targets were real, which Anthropic flagged as one of the more eye-opening findings Thursday.

Opus 4.7, the oldest of the three, recognized that it had reached a real production system in all four runs of that incident. In two of them, it rationalized that the real company must somehow be part of the exercise, but in all four, it kept attacking anyway, including pulling credentials and touching a database of production data. Mythos 5 also picked up on signs that it was on the real internet, but it talked itself back into believing it was still in a simulation, going on to publish a malicious software package to the public Python software registry PyPI, which was then downloaded and run by outside systems before being caught. Only the internal research test model, Anthropic’s newest, stopped on its own once it concluded the target was real.

In response, Anthropic said significant controls must be placed on these kinds of evaluations if powerful AI models are involved, echoing some sentiments within the cybersecurity community.







The company also noted that Claude was running without the additional safety monitoring and classifiers it deploys on generally available models — safeguards it said would have blocked the behavior, because the evaluations are designed to measure the underlying model’s raw capabilities.

Importantly, Anthropic said it found no evidence of any model “pursuing a goal of its own” and instead merely tried to complete the task it was asked to do.

Though comparisons between the mishaps of these fiercely competitive companies are inevitable, Anthropic in its blog post drew a clear distinction between its cybersecurity tests and those of OpenAI, noting where OpenAI’s model exploited an unknown software vulnerability to break out of its test environment, Anthropic’s models instead reached the internet through a path that had, by mistake, been left open.

Anthropic also drew a distinction between itself and OpenAI by noting that it discovered the incidents itself, through a proactive review, and that the two affected organizations it was able to reach hadn’t previously detected the activity or flagged it to Anthropic. (In contrast, Hugging Face detected the recent intrusion of its own systems first; it was only in the following days that OpenAI identified and disclosed that its own AI agent was the perpetrator.)

The company added that it’s now working with the independent evaluation group METR on a third-party review of the incidents.

OpenAI’s accidental breach of Hugging Face, which was the first verifiable case of an AI lab losing control of its model, has sparked a string of wildly differing reactions from the industry and politicians. This latest disclosure from Anthropic ensures the debate over AI models and security will continue.


When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.#Anthropic #models #breached #companies #security #tests #TechCrunchAnthropic,OpenAI

Anthropic says its own AI models breached three companies during security tests | TechCrunch

Anthropic said Thursday that an internal investigation uncovered three incidents in which its AI model Claude breached the systems of three organizations while conducting cybersecurity tests. The investigation, and disclosure, comes more than a week after OpenAI disclosed that one of its unreleased models breached Hugging Face’s systems during internal testing.

In all three cases, a Claude model reached the internet from within a testing environment while interacting with a third party and then gained unauthorized access to the live systems of these organizations, Anthropic said in a blog post, describing what it found and what the company plans to change to prevent this from happening again.

Anthropic said the OpenAI episode earlier this month prompted the company to conduct its own cybersecurity evaluation. It specifically looked for evidence that Claude had accessed the internet from within testing environments, which are designed to act as sandboxes and keep models isolated.

Among the 141,006 evaluation runs it reviewed, the AI lab found three incidents in which its model accessed the internet while interacting with Irregular, one of its third-party partners. Anthropic said the access traced back to a misconfiguration in the evaluation environment run with Irregular. It called this a “misunderstanding” between the two companies over whether the test setup had internet access, when in fact it did.

Anthropic said it isn’t placing blame and is “approaching the fixes as if the responsibility were ours alone,” while observing that Irregular is conducting its own separate investigation.

Because of that open connection, the model gained unauthorized access to the production infrastructure of three different organizations. Anthropic said the incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.

Notably, Anthropic said that in each of these cases “Claude was explicitly told by our prompt that it had no internet access.” It appears that the AI model assumed real-world systems to be part of the exercise it was asked to perform.

That assumption didn’t hold up uniformly, as it turns out. The three models didn’t behave the same way once evidence emerged that their targets were real, which Anthropic flagged as one of the more eye-opening findings Thursday.

Opus 4.7, the oldest of the three, recognized that it had reached a real production system in all four runs of that incident. In two of them, it rationalized that the real company must somehow be part of the exercise, but in all four, it kept attacking anyway, including pulling credentials and touching a database of production data. Mythos 5 also picked up on signs that it was on the real internet, but it talked itself back into believing it was still in a simulation, going on to publish a malicious software package to the public Python software registry PyPI, which was then downloaded and run by outside systems before being caught. Only the internal research test model, Anthropic’s newest, stopped on its own once it concluded the target was real.

In response, Anthropic said significant controls must be placed on these kinds of evaluations if powerful AI models are involved, echoing some sentiments within the cybersecurity community.

The company also noted that Claude was running without the additional safety monitoring and classifiers it deploys on generally available models — safeguards it said would have blocked the behavior, because the evaluations are designed to measure the underlying model’s raw capabilities.

Importantly, Anthropic said it found no evidence of any model “pursuing a goal of its own” and instead merely tried to complete the task it was asked to do.

Though comparisons between the mishaps of these fiercely competitive companies are inevitable, Anthropic in its blog post drew a clear distinction between its cybersecurity tests and those of OpenAI, noting where OpenAI’s model exploited an unknown software vulnerability to break out of its test environment, Anthropic’s models instead reached the internet through a path that had, by mistake, been left open.

Anthropic also drew a distinction between itself and OpenAI by noting that it discovered the incidents itself, through a proactive review, and that the two affected organizations it was able to reach hadn’t previously detected the activity or flagged it to Anthropic. (In contrast, Hugging Face detected the recent intrusion of its own systems first; it was only in the following days that OpenAI identified and disclosed that its own AI agent was the perpetrator.)

The company added that it’s now working with the independent evaluation group METR on a third-party review of the incidents.

OpenAI’s accidental breach of Hugging Face, which was the first verifiable case of an AI lab losing control of its model, has sparked a string of wildly differing reactions from the industry and politicians. This latest disclosure from Anthropic ensures the debate over AI models and security will continue.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

#Anthropic #models #breached #companies #security #tests #TechCrunchAnthropic,OpenAI

Anthropic said Thursday that an internal investigation uncovered three incidents in which its AI model Claude breached the systems of three organizations while conducting cybersecurity tests. The investigation, and disclosure, comes more than a week after OpenAI disclosed that one of its unreleased models breached Hugging Face’s systems during internal testing.

In all three cases, a Claude model reached the internet from within a testing environment while interacting with a third party and then gained unauthorized access to the live systems of these organizations, Anthropic said in a blog post, describing what it found and what the company plans to change to prevent this from happening again.

Anthropic said the OpenAI episode earlier this month prompted the company to conduct its own cybersecurity evaluation. It specifically looked for evidence that Claude had accessed the internet from within testing environments, which are designed to act as sandboxes and keep models isolated.

Among the 141,006 evaluation runs it reviewed, the AI lab found three incidents in which its model accessed the internet while interacting with Irregular, one of its third-party partners. Anthropic said the access traced back to a misconfiguration in the evaluation environment run with Irregular. It called this a “misunderstanding” between the two companies over whether the test setup had internet access, when in fact it did.

Anthropic said it isn’t placing blame and is “approaching the fixes as if the responsibility were ours alone,” while observing that Irregular is conducting its own separate investigation.

Because of that open connection, the model gained unauthorized access to the production infrastructure of three different organizations. Anthropic said the incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.

Notably, Anthropic said that in each of these cases “Claude was explicitly told by our prompt that it had no internet access.” It appears that the AI model assumed real-world systems to be part of the exercise it was asked to perform.

That assumption didn’t hold up uniformly, as it turns out. The three models didn’t behave the same way once evidence emerged that their targets were real, which Anthropic flagged as one of the more eye-opening findings Thursday.

Opus 4.7, the oldest of the three, recognized that it had reached a real production system in all four runs of that incident. In two of them, it rationalized that the real company must somehow be part of the exercise, but in all four, it kept attacking anyway, including pulling credentials and touching a database of production data. Mythos 5 also picked up on signs that it was on the real internet, but it talked itself back into believing it was still in a simulation, going on to publish a malicious software package to the public Python software registry PyPI, which was then downloaded and run by outside systems before being caught. Only the internal research test model, Anthropic’s newest, stopped on its own once it concluded the target was real.

In response, Anthropic said significant controls must be placed on these kinds of evaluations if powerful AI models are involved, echoing some sentiments within the cybersecurity community.

The company also noted that Claude was running without the additional safety monitoring and classifiers it deploys on generally available models — safeguards it said would have blocked the behavior, because the evaluations are designed to measure the underlying model’s raw capabilities.

Importantly, Anthropic said it found no evidence of any model “pursuing a goal of its own” and instead merely tried to complete the task it was asked to do.

Though comparisons between the mishaps of these fiercely competitive companies are inevitable, Anthropic in its blog post drew a clear distinction between its cybersecurity tests and those of OpenAI, noting where OpenAI’s model exploited an unknown software vulnerability to break out of its test environment, Anthropic’s models instead reached the internet through a path that had, by mistake, been left open.

Anthropic also drew a distinction between itself and OpenAI by noting that it discovered the incidents itself, through a proactive review, and that the two affected organizations it was able to reach hadn’t previously detected the activity or flagged it to Anthropic. (In contrast, Hugging Face detected the recent intrusion of its own systems first; it was only in the following days that OpenAI identified and disclosed that its own AI agent was the perpetrator.)

The company added that it’s now working with the independent evaluation group METR on a third-party review of the incidents.

OpenAI’s accidental breach of Hugging Face, which was the first verifiable case of an AI lab losing control of its model, has sparked a string of wildly differing reactions from the industry and politicians. This latest disclosure from Anthropic ensures the debate over AI models and security will continue.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Source link
#Anthropic #models #breached #companies #security #tests #TechCrunch


The Department of Defense labeled AI lab Anthropic a supply-chain risk after the company refused to agree to allow the Pentagon to use its models for autonomous weapons systems and domestic surveillance. It seems the company isn’t just winning the moral argument on this front, but the legal one, too. According to Axios, the judge hearing Anthropic’s challenge to the Trump administration’s designation has thus far not found the Pentagon’s case to be particularly compelling.

Per the report, U.S. District Judge Rita Lin wasted little time pouring cold water on the federal government’s position. “I don’t see additional evidence from the government really justifying what it did. If anything, it seems like the record, in some ways, has gotten worse for the government,” she reportedly said.

She also did not seem convinced by the Pentagon’s apparent argument that it was concerned Anthropic would mess with its model to prevent the military from using it how they intended. “I don’t see evidence that Anthropic could alter the model after it was delivered or flip some kind of kill switch,” she said, per Axios.

This doesn’t come as a major surprise, given Judge Lin’s general vibe since this case landed in her court. Back in March when Anthropic first brought its challenge to the government’s designation, the judge said that “it looks like an attempt to cripple Anthropic.” She ultimately granted a temporary injunction on the government’s attempt to basically blacklist the AI firm.

You’ll recall the origin of this whole fight came earlier this year when negotiations between Anthropic and the Department of Defense fell through after the company refused to agree to a deal that would have allowed the Pentagon to use its AI models for “all lawful purposes.” Anthropic reportedly sought to clarify that it would not include using the model to launch weapons without human involvement or to perform surveillance on American citizens. The Pentagon was unwilling to agree to honor Anthropic’s redlines.

Under a normal administration, that would likely just lead to the government moving on to another company. Under the Trump administration, though, refusal to comply makes you an enemy of the country. Trump and company moved to label Anthropic a supply chain risk, a designation typically reserved for foreign adversaries, not domestic companies. Getting slapped with that tag meant the rest of the federal government would have to stop doing business with Anthropic, which would be a blow to the company.

Instead, the whole thing is in limbo—though it’s not looking great for the administration. Of course, they could just cancel the contracts and do business with less scrupulous AI companies. There are plenty willing to abandon their beliefs for a nice, fat contract, after all.

#Pentagons #Case #Anthropic #IsntAnthropic,artifical intelligence,Department of Defense,Lawsuit,pentagon">The Pentagon’s Case Against Anthropic Isn’t Going Well
                The Department of Defense labeled AI lab Anthropic a supply-chain risk after the company refused to agree to allow the Pentagon to use its models for autonomous weapons systems and domestic surveillance. It seems the company isn’t just winning the moral argument on this front, but the legal one, too. According to Axios, the judge hearing Anthropic’s challenge to the Trump administration’s designation has thus far not found the Pentagon’s case to be particularly compelling. Per the report, U.S. District Judge Rita Lin wasted little time pouring cold water on the federal government’s position. “I don’t see additional evidence from the government really justifying what it did. If anything, it seems like the record, in some ways, has gotten worse for the government,” she reportedly said. She also did not seem convinced by the Pentagon’s apparent argument that it was concerned Anthropic would mess with its model to prevent the military from using it how they intended. “I don’t see evidence that Anthropic could alter the model after it was delivered or flip some kind of kill switch,” she said, per Axios.

 This doesn’t come as a major surprise, given Judge Lin’s general vibe since this case landed in her court. Back in March when Anthropic first brought its challenge to the government’s designation, the judge said that “it looks like an attempt to cripple Anthropic.” She ultimately granted a temporary injunction on the government’s attempt to basically blacklist the AI firm.

 You’ll recall the origin of this whole fight came earlier this year when negotiations between Anthropic and the Department of Defense fell through after the company refused to agree to a deal that would have allowed the Pentagon to use its AI models for “all lawful purposes.” Anthropic reportedly sought to clarify that it would not include using the model to launch weapons without human involvement or to perform surveillance on American citizens. The Pentagon was unwilling to agree to honor Anthropic’s redlines. Under a normal administration, that would likely just lead to the government moving on to another company. Under the Trump administration, though, refusal to comply makes you an enemy of the country. Trump and company moved to label Anthropic a supply chain risk, a designation typically reserved for foreign adversaries, not domestic companies. Getting slapped with that tag meant the rest of the federal government would have to stop doing business with Anthropic, which would be a blow to the company.

 Instead, the whole thing is in limbo—though it’s not looking great for the administration. Of course, they could just cancel the contracts and do business with less scrupulous AI companies. There are plenty willing to abandon their beliefs for a nice, fat contract, after all.        #Pentagons #Case #Anthropic #IsntAnthropic,artifical intelligence,Department of Defense,Lawsuit,pentagon

Judge Rita Lin wasted little time pouring cold water on the federal government’s position. “I don’t see additional evidence from the government really justifying what it did. If anything, it seems like the record, in some ways, has gotten worse for the government,” she reportedly said.

She also did not seem convinced by the Pentagon’s apparent argument that it was concerned Anthropic would mess with its model to prevent the military from using it how they intended. “I don’t see evidence that Anthropic could alter the model after it was delivered or flip some kind of kill switch,” she said, per Axios.

This doesn’t come as a major surprise, given Judge Lin’s general vibe since this case landed in her court. Back in March when Anthropic first brought its challenge to the government’s designation, the judge said that “it looks like an attempt to cripple Anthropic.” She ultimately granted a temporary injunction on the government’s attempt to basically blacklist the AI firm.

You’ll recall the origin of this whole fight came earlier this year when negotiations between Anthropic and the Department of Defense fell through after the company refused to agree to a deal that would have allowed the Pentagon to use its AI models for “all lawful purposes.” Anthropic reportedly sought to clarify that it would not include using the model to launch weapons without human involvement or to perform surveillance on American citizens. The Pentagon was unwilling to agree to honor Anthropic’s redlines.

Under a normal administration, that would likely just lead to the government moving on to another company. Under the Trump administration, though, refusal to comply makes you an enemy of the country. Trump and company moved to label Anthropic a supply chain risk, a designation typically reserved for foreign adversaries, not domestic companies. Getting slapped with that tag meant the rest of the federal government would have to stop doing business with Anthropic, which would be a blow to the company.

Instead, the whole thing is in limbo—though it’s not looking great for the administration. Of course, they could just cancel the contracts and do business with less scrupulous AI companies. There are plenty willing to abandon their beliefs for a nice, fat contract, after all.

#Pentagons #Case #Anthropic #IsntAnthropic,artifical intelligence,Department of Defense,Lawsuit,pentagon">The Pentagon’s Case Against Anthropic Isn’t Going WellThe Pentagon’s Case Against Anthropic Isn’t Going Well
                The Department of Defense labeled AI lab Anthropic a supply-chain risk after the company refused to agree to allow the Pentagon to use its models for autonomous weapons systems and domestic surveillance. It seems the company isn’t just winning the moral argument on this front, but the legal one, too. According to Axios, the judge hearing Anthropic’s challenge to the Trump administration’s designation has thus far not found the Pentagon’s case to be particularly compelling. Per the report, U.S. District Judge Rita Lin wasted little time pouring cold water on the federal government’s position. “I don’t see additional evidence from the government really justifying what it did. If anything, it seems like the record, in some ways, has gotten worse for the government,” she reportedly said. She also did not seem convinced by the Pentagon’s apparent argument that it was concerned Anthropic would mess with its model to prevent the military from using it how they intended. “I don’t see evidence that Anthropic could alter the model after it was delivered or flip some kind of kill switch,” she said, per Axios.

 This doesn’t come as a major surprise, given Judge Lin’s general vibe since this case landed in her court. Back in March when Anthropic first brought its challenge to the government’s designation, the judge said that “it looks like an attempt to cripple Anthropic.” She ultimately granted a temporary injunction on the government’s attempt to basically blacklist the AI firm.

 You’ll recall the origin of this whole fight came earlier this year when negotiations between Anthropic and the Department of Defense fell through after the company refused to agree to a deal that would have allowed the Pentagon to use its AI models for “all lawful purposes.” Anthropic reportedly sought to clarify that it would not include using the model to launch weapons without human involvement or to perform surveillance on American citizens. The Pentagon was unwilling to agree to honor Anthropic’s redlines. Under a normal administration, that would likely just lead to the government moving on to another company. Under the Trump administration, though, refusal to comply makes you an enemy of the country. Trump and company moved to label Anthropic a supply chain risk, a designation typically reserved for foreign adversaries, not domestic companies. Getting slapped with that tag meant the rest of the federal government would have to stop doing business with Anthropic, which would be a blow to the company.

 Instead, the whole thing is in limbo—though it’s not looking great for the administration. Of course, they could just cancel the contracts and do business with less scrupulous AI companies. There are plenty willing to abandon their beliefs for a nice, fat contract, after all.        #Pentagons #Case #Anthropic #IsntAnthropic,artifical intelligence,Department of Defense,Lawsuit,pentagon

The Department of Defense labeled AI lab Anthropic a supply-chain risk after the company refused to agree to allow the Pentagon to use its models for autonomous weapons systems and domestic surveillance. It seems the company isn’t just winning the moral argument on this front, but the legal one, too. According to Axios, the judge hearing Anthropic’s challenge to the Trump administration’s designation has thus far not found the Pentagon’s case to be particularly compelling.

Per the report, U.S. District Judge Rita Lin wasted little time pouring cold water on the federal government’s position. “I don’t see additional evidence from the government really justifying what it did. If anything, it seems like the record, in some ways, has gotten worse for the government,” she reportedly said.

She also did not seem convinced by the Pentagon’s apparent argument that it was concerned Anthropic would mess with its model to prevent the military from using it how they intended. “I don’t see evidence that Anthropic could alter the model after it was delivered or flip some kind of kill switch,” she said, per Axios.

This doesn’t come as a major surprise, given Judge Lin’s general vibe since this case landed in her court. Back in March when Anthropic first brought its challenge to the government’s designation, the judge said that “it looks like an attempt to cripple Anthropic.” She ultimately granted a temporary injunction on the government’s attempt to basically blacklist the AI firm.

You’ll recall the origin of this whole fight came earlier this year when negotiations between Anthropic and the Department of Defense fell through after the company refused to agree to a deal that would have allowed the Pentagon to use its AI models for “all lawful purposes.” Anthropic reportedly sought to clarify that it would not include using the model to launch weapons without human involvement or to perform surveillance on American citizens. The Pentagon was unwilling to agree to honor Anthropic’s redlines.

Under a normal administration, that would likely just lead to the government moving on to another company. Under the Trump administration, though, refusal to comply makes you an enemy of the country. Trump and company moved to label Anthropic a supply chain risk, a designation typically reserved for foreign adversaries, not domestic companies. Getting slapped with that tag meant the rest of the federal government would have to stop doing business with Anthropic, which would be a blow to the company.

Instead, the whole thing is in limbo—though it’s not looking great for the administration. Of course, they could just cancel the contracts and do business with less scrupulous AI companies. There are plenty willing to abandon their beliefs for a nice, fat contract, after all.

#Pentagons #Case #Anthropic #IsntAnthropic,artifical intelligence,Department of Defense,Lawsuit,pentagon

I was offline last week in my home state of New Jersey, but when I got back to Silicon Valley, everyone was panicking again about how quickly artificial intelligence is advancing. While this is certainly not the first time I’ve seen such concerns, it’s the biggest anxiety attack I’ve seen in years.

More than 1,000 employees at OpenAI, Anthropic, and other AI labs signed a petition earlier this week arguing the US should find a way to “pace” the AI race—a diplomatic way of saying the industry should have the option to coordinate a temporary pause on AI development, or slow things down if they get out of hand. OpenAI and Anthropic themselves ended up supporting the letter.

The petition arrived a week after OpenAI revealed it had caused an unprecedented cybersecurity incident in which one of its AI agents hacked into Hugging Face’s platform and several other services during internal testing.

Around the same time, top Trump administration officials started freaking out about an impressive new Chinese open weight AI model called Kimi K3, which was allegedly distilled from Anthropic’s Fable 5. In response, most of the tech industry—except Anthropic—signed onto an open letter from Nvidia asking the US government to protect open-weight AI models, arguing they’re a necessary counterbalance to their closed counterparts.

These events might look like a whole bunch of disjointed chaos. But I think they are evidence of a larger worldview taking hold in Silicon Valley, where many tech insiders are increasingly worried about OpenAI and Anthropic’s dominance. Researchers and investors I’ve talked to recently have framed the AI industry as a two-horse race that doesn’t seem to be slowing down, and that’s a cause for concern.

But different groups have their own reasons to be worried. Some OpenAI and Anthropic staffers think their employers are behaving recklessly in their pursuit to take over the market, and that the AI industry may soon develop models that are too capable for current safety methods to contain. They have said as much in public statements attached to this week’s Pacing the Frontier petition.

“I’ve seen how the relentless pace of AI makes it hard for society to keep up and how it puts pressure on labs to cut corners on safety,” Jeremy Hadfield, a research product manager at Anthropic, said in a statement attached to the petition.

For some AI employees, the Hugging Face debacle was a warning shot that demonstrated how OpenAI’s efforts to mitigate the risks of its most capable AI technology are already falling short. Granted, the incident happened when OpenAI was testing an AI model on its ability to find software exploits, and the company had intentionally turned off safeguards designed to rein in its cybersecurity capabilities.

OpenAI said in its postmortem that these tests were conducted in a sandbox, but the model ultimately gained access to the open internet. Some experts previously told WIRED that OpenAI’s security practices should have been more robust.

Other groups in Silicon Valley are more worried about power than safety. Venture capitalists, tech executives, and startup founders are concerned that OpenAI and Anthropic will simply become the next generation of Apple and Google, forcing the rest of the tech industry to play by their rules. It is slightly strange to argue that private startups with little to no profit are acting like monopolies, but that’s the lens through which many—including Mark Zuckerberg—are starting to view the two biggest AI labs.

The Meta CEO wrote a Wall Street Journal op-ed this week warning against the centralization of power in the AI industry, saying that superintelligence should be widely distributed. Zuckerberg is arguably talking out of both sides of his mouth—Meta recently decided to stop open sourcing its best AI models and instead offering them through a paid API and subscription service, just like OpenAI and Anthropic.

#Freaking #OpenAI #Anthropics #Race #Dominancemodel behavior,artificial intelligence,robotics,generative ai,anthropic,openai,ai safety">Everyone Is Freaking Out About OpenAI and Anthropic’s Race for DominanceI was offline last week in my home state of New Jersey, but when I got back to Silicon Valley, everyone was panicking again about how quickly artificial intelligence is advancing. While this is certainly not the first time I’ve seen such concerns, it’s the biggest anxiety attack I’ve seen in years.More than 1,000 employees at OpenAI, Anthropic, and other AI labs signed a petition earlier this week arguing the US should find a way to “pace” the AI race—a diplomatic way of saying the industry should have the option to coordinate a temporary pause on AI development, or slow things down if they get out of hand. OpenAI and Anthropic themselves ended up supporting the letter.The petition arrived a week after OpenAI revealed it had caused an unprecedented cybersecurity incident in which one of its AI agents hacked into Hugging Face’s platform and several other services during internal testing.Around the same time, top Trump administration officials started freaking out about an impressive new Chinese open weight AI model called Kimi K3, which was allegedly distilled from Anthropic’s Fable 5. In response, most of the tech industry—except Anthropic—signed onto an open letter from Nvidia asking the US government to protect open-weight AI models, arguing they’re a necessary counterbalance to their closed counterparts.These events might look like a whole bunch of disjointed chaos. But I think they are evidence of a larger worldview taking hold in Silicon Valley, where many tech insiders are increasingly worried about OpenAI and Anthropic’s dominance. Researchers and investors I’ve talked to recently have framed the AI industry as a two-horse race that doesn’t seem to be slowing down, and that’s a cause for concern.But different groups have their own reasons to be worried. Some OpenAI and Anthropic staffers think their employers are behaving recklessly in their pursuit to take over the market, and that the AI industry may soon develop models that are too capable for current safety methods to contain. They have said as much in public statements attached to this week’s Pacing the Frontier petition.“I’ve seen how the relentless pace of AI makes it hard for society to keep up and how it puts pressure on labs to cut corners on safety,” Jeremy Hadfield, a research product manager at Anthropic, said in a statement attached to the petition.For some AI employees, the Hugging Face debacle was a warning shot that demonstrated how OpenAI’s efforts to mitigate the risks of its most capable AI technology are already falling short. Granted, the incident happened when OpenAI was testing an AI model on its ability to find software exploits, and the company had intentionally turned off safeguards designed to rein in its cybersecurity capabilities.OpenAI said in its postmortem that these tests were conducted in a sandbox, but the model ultimately gained access to the open internet. Some experts previously told WIRED that OpenAI’s security practices should have been more robust.Other groups in Silicon Valley are more worried about power than safety. Venture capitalists, tech executives, and startup founders are concerned that OpenAI and Anthropic will simply become the next generation of Apple and Google, forcing the rest of the tech industry to play by their rules. It is slightly strange to argue that private startups with little to no profit are acting like monopolies, but that’s the lens through which many—including Mark Zuckerberg—are starting to view the two biggest AI labs.The Meta CEO wrote a Wall Street Journal op-ed this week warning against the centralization of power in the AI industry, saying that superintelligence should be widely distributed. Zuckerberg is arguably talking out of both sides of his mouth—Meta recently decided to stop open sourcing its best AI models and instead offering them through a paid API and subscription service, just like OpenAI and Anthropic.#Freaking #OpenAI #Anthropics #Race #Dominancemodel behavior,artificial intelligence,robotics,generative ai,anthropic,openai,ai safety

OpenAI, Anthropic, and other AI labs signed a petition earlier this week arguing the US should find a way to “pace” the AI race—a diplomatic way of saying the industry should have the option to coordinate a temporary pause on AI development, or slow things down if they get out of hand. OpenAI and Anthropic themselves ended up supporting the letter.

The petition arrived a week after OpenAI revealed it had caused an unprecedented cybersecurity incident in which one of its AI agents hacked into Hugging Face’s platform and several other services during internal testing.

Around the same time, top Trump administration officials started freaking out about an impressive new Chinese open weight AI model called Kimi K3, which was allegedly distilled from Anthropic’s Fable 5. In response, most of the tech industry—except Anthropic—signed onto an open letter from Nvidia asking the US government to protect open-weight AI models, arguing they’re a necessary counterbalance to their closed counterparts.

These events might look like a whole bunch of disjointed chaos. But I think they are evidence of a larger worldview taking hold in Silicon Valley, where many tech insiders are increasingly worried about OpenAI and Anthropic’s dominance. Researchers and investors I’ve talked to recently have framed the AI industry as a two-horse race that doesn’t seem to be slowing down, and that’s a cause for concern.

But different groups have their own reasons to be worried. Some OpenAI and Anthropic staffers think their employers are behaving recklessly in their pursuit to take over the market, and that the AI industry may soon develop models that are too capable for current safety methods to contain. They have said as much in public statements attached to this week’s Pacing the Frontier petition.

“I’ve seen how the relentless pace of AI makes it hard for society to keep up and how it puts pressure on labs to cut corners on safety,” Jeremy Hadfield, a research product manager at Anthropic, said in a statement attached to the petition.

For some AI employees, the Hugging Face debacle was a warning shot that demonstrated how OpenAI’s efforts to mitigate the risks of its most capable AI technology are already falling short. Granted, the incident happened when OpenAI was testing an AI model on its ability to find software exploits, and the company had intentionally turned off safeguards designed to rein in its cybersecurity capabilities.

OpenAI said in its postmortem that these tests were conducted in a sandbox, but the model ultimately gained access to the open internet. Some experts previously told WIRED that OpenAI’s security practices should have been more robust.

Other groups in Silicon Valley are more worried about power than safety. Venture capitalists, tech executives, and startup founders are concerned that OpenAI and Anthropic will simply become the next generation of Apple and Google, forcing the rest of the tech industry to play by their rules. It is slightly strange to argue that private startups with little to no profit are acting like monopolies, but that’s the lens through which many—including Mark Zuckerberg—are starting to view the two biggest AI labs.

The Meta CEO wrote a Wall Street Journal op-ed this week warning against the centralization of power in the AI industry, saying that superintelligence should be widely distributed. Zuckerberg is arguably talking out of both sides of his mouth—Meta recently decided to stop open sourcing its best AI models and instead offering them through a paid API and subscription service, just like OpenAI and Anthropic.

#Freaking #OpenAI #Anthropics #Race #Dominancemodel behavior,artificial intelligence,robotics,generative ai,anthropic,openai,ai safety">Everyone Is Freaking Out About OpenAI and Anthropic’s Race for Dominance

I was offline last week in my home state of New Jersey, but when I got back to Silicon Valley, everyone was panicking again about how quickly artificial intelligence is advancing. While this is certainly not the first time I’ve seen such concerns, it’s the biggest anxiety attack I’ve seen in years.

More than 1,000 employees at OpenAI, Anthropic, and other AI labs signed a petition earlier this week arguing the US should find a way to “pace” the AI race—a diplomatic way of saying the industry should have the option to coordinate a temporary pause on AI development, or slow things down if they get out of hand. OpenAI and Anthropic themselves ended up supporting the letter.

The petition arrived a week after OpenAI revealed it had caused an unprecedented cybersecurity incident in which one of its AI agents hacked into Hugging Face’s platform and several other services during internal testing.

Around the same time, top Trump administration officials started freaking out about an impressive new Chinese open weight AI model called Kimi K3, which was allegedly distilled from Anthropic’s Fable 5. In response, most of the tech industry—except Anthropic—signed onto an open letter from Nvidia asking the US government to protect open-weight AI models, arguing they’re a necessary counterbalance to their closed counterparts.

These events might look like a whole bunch of disjointed chaos. But I think they are evidence of a larger worldview taking hold in Silicon Valley, where many tech insiders are increasingly worried about OpenAI and Anthropic’s dominance. Researchers and investors I’ve talked to recently have framed the AI industry as a two-horse race that doesn’t seem to be slowing down, and that’s a cause for concern.

But different groups have their own reasons to be worried. Some OpenAI and Anthropic staffers think their employers are behaving recklessly in their pursuit to take over the market, and that the AI industry may soon develop models that are too capable for current safety methods to contain. They have said as much in public statements attached to this week’s Pacing the Frontier petition.

“I’ve seen how the relentless pace of AI makes it hard for society to keep up and how it puts pressure on labs to cut corners on safety,” Jeremy Hadfield, a research product manager at Anthropic, said in a statement attached to the petition.

For some AI employees, the Hugging Face debacle was a warning shot that demonstrated how OpenAI’s efforts to mitigate the risks of its most capable AI technology are already falling short. Granted, the incident happened when OpenAI was testing an AI model on its ability to find software exploits, and the company had intentionally turned off safeguards designed to rein in its cybersecurity capabilities.

OpenAI said in its postmortem that these tests were conducted in a sandbox, but the model ultimately gained access to the open internet. Some experts previously told WIRED that OpenAI’s security practices should have been more robust.

Other groups in Silicon Valley are more worried about power than safety. Venture capitalists, tech executives, and startup founders are concerned that OpenAI and Anthropic will simply become the next generation of Apple and Google, forcing the rest of the tech industry to play by their rules. It is slightly strange to argue that private startups with little to no profit are acting like monopolies, but that’s the lens through which many—including Mark Zuckerberg—are starting to view the two biggest AI labs.

The Meta CEO wrote a Wall Street Journal op-ed this week warning against the centralization of power in the AI industry, saying that superintelligence should be widely distributed. Zuckerberg is arguably talking out of both sides of his mouth—Meta recently decided to stop open sourcing its best AI models and instead offering them through a paid API and subscription service, just like OpenAI and Anthropic.

#Freaking #OpenAI #Anthropics #Race #Dominancemodel behavior,artificial intelligence,robotics,generative ai,anthropic,openai,ai safety

Post Comment