How AI guardrails are impeding the work of offensive cybersecurity researchers | TechCrunch
For months, AI giants have devised special vetted programs and strict guardrails to limit the use of their models by malicious hackers. But these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers.
In June, the U.S. government slapped export control restrictions on Anthropic’s much-hyped AI models Mythos and Fable. The move was prompted at least in part by a report that claimed it was possible to bypass the models’ guardrails designed to prevent users from using them to build and execute malicious cyberattacks.
Regardless of whether the incident was really motivated by fears of a jailbreak, the fact is that Anthropic has repeatedly marketed Mythos as some kind of doomsday cybermachine that can only be given to carefully vetted users, and even then with strict guardrails in place. (The export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government’s review process.)
That kind of gatekeeping isn’t unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs they can apply to get vetted and — if approved — access models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification Program.
These guardrails have been widely criticized, particularly by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.
During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said that, “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.”
Dowd has spent decades finding and selling “zero-days” — previously unknown software flaws and the exploits that take advantage of them — to Western governments, rather than reporting them to the software makers so they get patched. Governments pay a premium for vulnerabilities precisely because they stay open, which is useful for intelligence operations.
Dowd admitted his work may make him biased, but he isn’t alone. Several people who work in offensive cybersecurity — they proactively probe systems for weaknesses — described to TechCrunch how they use AI tools and deal with their guardrails.
Chris Anley, the chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders, he said.
“This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” said Anley. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”
It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.”
When he and his colleagues run into such a roadblock, they sometimes fall back on open source AI models that come with no guardrails at all.
Paolo Stagno, the chief technology officer at Crowdfense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying AI companies “essentially treat customers like children who need babysitting” with their vetted programs and guardrails.
Stagno said he and his colleagues do use frontier models — but only for reverse engineering. They avoid using AI to help find vulnerabilities or build exploits, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, he said, they use open source models run locally, as they do not rely on sharing data outside of the model.
Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails are not impeding his work. That’s because he doesn’t use AI for offensive work; instead, he uses it for initial reverse engineering, to understand the code he’s analyzing, and to build supporting tools. For that, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities.
“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” said Cali. “I am jealous of my bugs, and I like this game too much to let models play it for me.”
One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity because he isn’t authorized to talk to the press, said his employer isn’t part of Anthropic’s CVP program and as a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict.
“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person said.
Chris Thompson — chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive security and AI-focused event — said that in his experience using the frontier AI models, the guardrails can be inconsistent and work differently every day. That’s true even inside the looser boundaries of Anthropic’s and OpenAI’s vetted programs.
“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” said Thompson. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”
Consequently, researchers rely on or get pushed toward Chinese open source models like GLM — freely downloadable models that can be run locally with no vetting or usage restrictions — said Thompson.
“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. “I think it’s more harmful than good to have these guardrails in place.”
Rather than tightening restrictions further, Thompson called for the AI frontier labs to open up their programs, provide responsible access, and hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race.
“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” said Thompson. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”
For months, AI giants have devised special vetted programs and strict guardrails to limit the use of their models by malicious hackers. But these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers.
In June, the U.S. government slapped export control restrictions on Anthropic’s much-hyped AI models Mythos and Fable. The move was prompted at least in part by a report that claimed it was possible to bypass the models’ guardrails designed to prevent users from using them to build and execute malicious cyberattacks.
Regardless of whether the incident was really motivated by fears of a jailbreak, the fact is that Anthropic has repeatedly marketed Mythos as some kind of doomsday cybermachine that can only be given to carefully vetted users, and even then with strict guardrails in place. (The export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government’s review process.)
That kind of gatekeeping isn’t unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs they can apply to get vetted and — if approved — access models with fewer cybersecurity restrictions: OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification Program.
These guardrails have been widely criticized, particularly by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.
During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said that, “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.”
Dowd has spent decades finding and selling “zero-days” — previously unknown software flaws and the exploits that take advantage of them — to Western governments, rather than reporting them to the software makers so they get patched. Governments pay a premium for vulnerabilities precisely because they stay open, which is useful for intelligence operations.
Dowd admitted his work may make him biased, but he isn’t alone. Several people who work in offensive cybersecurity — they proactively probe systems for weaknesses — described to TechCrunch how they use AI tools and deal with their guardrails.
Chris Anley, the chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders, he said.
“This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” said Anley. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”
It’s “like a hammer,” he continued. “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.”
When he and his colleagues run into such a roadblock, they sometimes fall back on open source AI models that come with no guardrails at all.
Paolo Stagno, the chief technology officer at Crowdfense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying AI companies “essentially treat customers like children who need babysitting” with their vetted programs and guardrails.
Stagno said he and his colleagues do use frontier models — but only for reverse engineering. They avoid using AI to help find vulnerabilities or build exploits, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, he said, they use open source models run locally, as they do not rely on sharing data outside of the model.
Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails are not impeding his work. That’s because he doesn’t use AI for offensive work; instead, he uses it for initial reverse engineering, to understand the code he’s analyzing, and to build supporting tools. For that, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities.
“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” said Cali. “I am jealous of my bugs, and I like this game too much to let models play it for me.”
One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity because he isn’t authorized to talk to the press, said his employer isn’t part of Anthropic’s CVP program and as a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict.
“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the person said.
Chris Thompson — chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive security and AI-focused event — said that in his experience using the frontier AI models, the guardrails can be inconsistent and work differently every day. That’s true even inside the looser boundaries of Anthropic’s and OpenAI’s vetted programs.
“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” said Thompson. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”
Consequently, researchers rely on or get pushed toward Chinese open source models like GLM — freely downloadable models that can be run locally with no vetting or usage restrictions — said Thompson.
“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he said. “I think it’s more harmful than good to have these guardrails in place.”
Rather than tightening restrictions further, Thompson called for the AI frontier labs to open up their programs, provide responsible access, and hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race.
“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” said Thompson. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
#guardrails #impeding #work #offensive #cybersecurity #researchers #TechCrunchcybersecurity,Zero-days
![This former notorious red-light district is now one of the world’s top AI hubs | TechCrunch
What every U.K. AI startup wants to know these days is, how can I get office space in King’s Cross?
The area is so hot that a VC firm allegedly recently won a deal by promising a founder office space in the neighborhood. “We stop at nothing to win deals [for] and to support” founders, “including helping them source office space when needed,” the firm told me when asked about the rumor, declining to confirm or deny any details.
The neighborhood’s popularity began back in 2016 when DeepMind — then newly acquired by Google — moved in. Soon after, a flood of AI startups followed, wanting to be around the Google DeepMind magic. Today, they hope to take advantage of the cluster of AI talent there.
This has transformed King’s Cross into one of the world’s top AI hubs, rivaled only by San Francisco and Beijing. Around London, it’s known by the sobriquet “Knowledge Quarter,” as it’s home to names like OpenAI, Meta, Isomorphic Labs, Cusp AI, Wayne, Recursive, and, a little farther down the road, Synthesia and Anthropic. The European Technology Network (ETN) just moved into a glossy new office nearby, while University College London sits around the corner.
Mixed in with the new developments are trendy food spots like Hoppers and BAO. Hop a train from King’s Cross, and founders can be in Cambridge in 45 minutes to source talent or can be in Paris in two hours to strike a deal.
Who would have guessed that a little more than 20 years ago, this was one of the seediest areas in London?
“In the ’80s, crack and heroin made the area a major narcotics market,” Hussein Kanji, an investor at Hoxton Ventures, said, recalling syringes in tree trunks and gangs patrolling the streets. “In 1982, the local church was occupied by the English Collective of Prostitutes for 12 straight days.” Then, in the early 2000s, a real estate developer had a dream and, well, “now it is the AI hotbed of the United Kingdom,” Kanji said. “What a change.” Around 18 months ago, his portfolio company BioCorteX moved from the neighborhood Holborn to the Jellicoe building in King’s Cross, hoping to be near the action. “Lots going on in London right now,” Nik Sharma, co-founder of BioCorteX, told me. “Lots of hyperscalers moving in.” That includes, reportedly, Jeff Bezos’ AI company Prometheus, which is also said to be in talks to move into the Jellicoe.
There are around 3,600 AI startups in London, which, together, have raised around .1 billion out of the .8 billion raised in the city since late July, according to Dealroom. Since the start of June, AI-related startups have leased more than 1 million square feet of office space in London, according to the real estate firm Knight Frank. With that, prime rents in King’s Cross have risen 18% over the past three years, Chris Dunn, a commercial insight associate at the firm, told me. That percentage represents only the largest leases encompassing at least 10,000 square feet, like the ones OpenAI and Prometheus are signing. The shorter deals go for even more, he said, and now the vacancy rate for conventional office space is just 0.9%. “Demand has outstripped supply,” he continued.
Today, one of the big topics of the area is sovereignty. It was a wake-up call for many when Anthropic shut off access to Mythos and Fable this summer, leaving some in the ecosystem to conclude: “We’d better look after ourselves,” Saul Klein, co-founder of the VC firm Phoenix Court, told me.
Phoenix Court is located in the King’s Cross area and has three portfolio companies in the vicinity, including Olix (which just announced a .3 billion valuation), Early Health and CoMind. Robin Klein, co-founder of the firm, said the shutdown of Fable and Mythos access was a “small but sharp reminder that Europe can’t simply rent its AI capabilities and capacity; it needs to build and hold some of its own.” King’s Cross, he said, is where much of this building is actually happening.
“The bigger question,” he continued, “is whether the U.K. builds the infrastructure, compute, energy, capital, to make this self-reliance durable, rather than just hosting outposts of U.S. labs.”
Image Credits:Phoenix Court
Top founders want to stay
Simon Kohl, founder of Latent Labs, has offices in King’s Cross and San Francisco. The London office, at the moment, is growing faster, and he’s more bullish than ever on the ecosystem, he said. “The mood right now feels less like London trying to catch up and more like London becoming one of the default places to start a serious AI company,” he said. Look around and you are likely to see Wayve testing its autonomous cars. Founded in 2017 by co-founder Alex Kendall, the unicorn is one of London’s biggest success stories.
“Ten years ago, building a frontier AI company from London felt like an unusual choice,” Kendall told me. “Now it feels like an obvious one.” Wayve moved into King’s Cross in 2018 looking for a space that could double as a garage — “a rare combination in Central London,” Kendall said. He has watched the ecosystem mature around him — and it’s now evident that a startup can stay in London, raise serious capital, hire world-class AI talent, and remain globally competitive, he said. Down the street from Anthropic’s new 158,000-square-foot office is the AI agent builder Sierra and the AI video platform Synthesia.
Laura Gonzalez Florez, Synthesia’s chief of staff and head of people, says the company moved into its glossy new office building a year ago to accommodate its growing team. They were drawn to the area for the same reason as everyone else: “It’s very close to the airport … very close to where a lot of investors are,” she said.
Image Credits:Synthesia
Around two-thirds of Synthesia’s engineers are remote, Gonzalez Florez said, letting the company tap into an affordable, international, and diverse talent pool and helping it scale faster. “From London, we can hire and work, without any problem, people from anywhere, from Slovenia to Portugal,” she said.
Unsurprisingly, London’s AI boom is also causing a talent war.U.K. AI job postings have skyrocketed in the past few years, per data from PwC. When Anthropic announced it moved into town earlier this year, it listed, for example, a salary range of £260,000 to £630,000 for a machine learning research engineer when the average salary in London for the same role is around £102,000. Some founders in the U.K., like those in Silicon Valley, are being forced to raise more and bigger rounds to keep up.
“The real test is whether more globally significant AI companies are founded, funded, and scaled from the U.K., while continuing to attract the world’s best talent to build them here,” Zain Ali, founder of the King’s Cross-based AI legal firm Centuro, told me. “If that continues to happen, King’s Cross won’t just be an AI hub. It’ll become one of the U.K.’s most important strategic assets.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.#Thisformernotorious #redlight #districtis #nowone #worlds #top #hubs #TechCrunchUK This former notorious red-light district is now one of the world’s top AI hubs | TechCrunch
What every U.K. AI startup wants to know these days is, how can I get office space in King’s Cross?
The area is so hot that a VC firm allegedly recently won a deal by promising a founder office space in the neighborhood. “We stop at nothing to win deals [for] and to support” founders, “including helping them source office space when needed,” the firm told me when asked about the rumor, declining to confirm or deny any details.
The neighborhood’s popularity began back in 2016 when DeepMind — then newly acquired by Google — moved in. Soon after, a flood of AI startups followed, wanting to be around the Google DeepMind magic. Today, they hope to take advantage of the cluster of AI talent there.
This has transformed King’s Cross into one of the world’s top AI hubs, rivaled only by San Francisco and Beijing. Around London, it’s known by the sobriquet “Knowledge Quarter,” as it’s home to names like OpenAI, Meta, Isomorphic Labs, Cusp AI, Wayne, Recursive, and, a little farther down the road, Synthesia and Anthropic. The European Technology Network (ETN) just moved into a glossy new office nearby, while University College London sits around the corner.
Mixed in with the new developments are trendy food spots like Hoppers and BAO. Hop a train from King’s Cross, and founders can be in Cambridge in 45 minutes to source talent or can be in Paris in two hours to strike a deal.
Who would have guessed that a little more than 20 years ago, this was one of the seediest areas in London?
“In the ’80s, crack and heroin made the area a major narcotics market,” Hussein Kanji, an investor at Hoxton Ventures, said, recalling syringes in tree trunks and gangs patrolling the streets. “In 1982, the local church was occupied by the English Collective of Prostitutes for 12 straight days.” Then, in the early 2000s, a real estate developer had a dream and, well, “now it is the AI hotbed of the United Kingdom,” Kanji said. “What a change.” Around 18 months ago, his portfolio company BioCorteX moved from the neighborhood Holborn to the Jellicoe building in King’s Cross, hoping to be near the action. “Lots going on in London right now,” Nik Sharma, co-founder of BioCorteX, told me. “Lots of hyperscalers moving in.” That includes, reportedly, Jeff Bezos’ AI company Prometheus, which is also said to be in talks to move into the Jellicoe.
There are around 3,600 AI startups in London, which, together, have raised around .1 billion out of the .8 billion raised in the city since late July, according to Dealroom. Since the start of June, AI-related startups have leased more than 1 million square feet of office space in London, according to the real estate firm Knight Frank. With that, prime rents in King’s Cross have risen 18% over the past three years, Chris Dunn, a commercial insight associate at the firm, told me. That percentage represents only the largest leases encompassing at least 10,000 square feet, like the ones OpenAI and Prometheus are signing. The shorter deals go for even more, he said, and now the vacancy rate for conventional office space is just 0.9%. “Demand has outstripped supply,” he continued.
Today, one of the big topics of the area is sovereignty. It was a wake-up call for many when Anthropic shut off access to Mythos and Fable this summer, leaving some in the ecosystem to conclude: “We’d better look after ourselves,” Saul Klein, co-founder of the VC firm Phoenix Court, told me.
Phoenix Court is located in the King’s Cross area and has three portfolio companies in the vicinity, including Olix (which just announced a .3 billion valuation), Early Health and CoMind. Robin Klein, co-founder of the firm, said the shutdown of Fable and Mythos access was a “small but sharp reminder that Europe can’t simply rent its AI capabilities and capacity; it needs to build and hold some of its own.” King’s Cross, he said, is where much of this building is actually happening.
“The bigger question,” he continued, “is whether the U.K. builds the infrastructure, compute, energy, capital, to make this self-reliance durable, rather than just hosting outposts of U.S. labs.”
Image Credits:Phoenix Court
Top founders want to stay
Simon Kohl, founder of Latent Labs, has offices in King’s Cross and San Francisco. The London office, at the moment, is growing faster, and he’s more bullish than ever on the ecosystem, he said. “The mood right now feels less like London trying to catch up and more like London becoming one of the default places to start a serious AI company,” he said. Look around and you are likely to see Wayve testing its autonomous cars. Founded in 2017 by co-founder Alex Kendall, the unicorn is one of London’s biggest success stories.
“Ten years ago, building a frontier AI company from London felt like an unusual choice,” Kendall told me. “Now it feels like an obvious one.” Wayve moved into King’s Cross in 2018 looking for a space that could double as a garage — “a rare combination in Central London,” Kendall said. He has watched the ecosystem mature around him — and it’s now evident that a startup can stay in London, raise serious capital, hire world-class AI talent, and remain globally competitive, he said. Down the street from Anthropic’s new 158,000-square-foot office is the AI agent builder Sierra and the AI video platform Synthesia.
Laura Gonzalez Florez, Synthesia’s chief of staff and head of people, says the company moved into its glossy new office building a year ago to accommodate its growing team. They were drawn to the area for the same reason as everyone else: “It’s very close to the airport … very close to where a lot of investors are,” she said.
Image Credits:Synthesia
Around two-thirds of Synthesia’s engineers are remote, Gonzalez Florez said, letting the company tap into an affordable, international, and diverse talent pool and helping it scale faster. “From London, we can hire and work, without any problem, people from anywhere, from Slovenia to Portugal,” she said.
Unsurprisingly, London’s AI boom is also causing a talent war.U.K. AI job postings have skyrocketed in the past few years, per data from PwC. When Anthropic announced it moved into town earlier this year, it listed, for example, a salary range of £260,000 to £630,000 for a machine learning research engineer when the average salary in London for the same role is around £102,000. Some founders in the U.K., like those in Silicon Valley, are being forced to raise more and bigger rounds to keep up.
“The real test is whether more globally significant AI companies are founded, funded, and scaled from the U.K., while continuing to attract the world’s best talent to build them here,” Zain Ali, founder of the King’s Cross-based AI legal firm Centuro, told me. “If that continues to happen, King’s Cross won’t just be an AI hub. It’ll become one of the U.K.’s most important strategic assets.”
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.#Thisformernotorious #redlight #districtis #nowone #worlds #top #hubs #TechCrunchUK](https://techcrunch.com/wp-content/uploads/2026/08/DM9A2852.jpg?w=680)

Post Comment