READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

AI Companies' Cybersecurity Guardrails Are Blocking Legitimate Security Researchers, Not Just Hackers

AI Companies' Cybersecurity Guardrails Are Blocking Legitimate Security Researchers, Not Just Hackers
Anthropic and OpenAI built strict guardrails into their AI models to stop criminals from weaponizing them for cyberattacks. Security researchers say those same guardrails now block the exact defensive work the companies claim to want, forcing legitimate professionals through vetting programs to do basic vulnerability testing.

AI companies built guardrails to stop hackers from using chatbots to write malware. Now the people whose job is to find security holes before criminals do say those same guardrails are getting in their way.

Anthropic and OpenAI both run vetting programs for cybersecurity researchers, according to TechCrunch. OpenAI calls its version Trusted Access for Cyber. Anthropic calls its Cyber Verification Program. Apply, get approved, and you get a model with fewer restrictions on cybersecurity-related questions.

The catch is that professional offensive security researchers, the ones paid to break into systems on purpose so companies can patch the holes, say the standard guardrails interfere with routine work that has nothing to do with crime.

Chris Anley, chief scientist at security consulting firm NCC Group, told TechCrunch that asking an AI model to try to exploit a bug is often the only way to confirm the bug is real and worth fixing. When a guardrail makes the model refuse to answer, Anley said, it doesn't stop bad actors, it stops defenders. His point: "fix this code" is simultaneously a defensive request and, in the wrong hands, a blueprint for finding a critical flaw. The model can't tell the difference between a researcher testing a client's system and a criminal probing a target.

The friction got a real-world test in June, when the U.S. government imposed export control restrictions on two Anthropic models, Mythos and Fable, according to TechCrunch. The restrictions followed a report claiming it was possible to bypass the guardrails meant to stop the models from being used to build and run cyberattacks.

Those controls didn't last. Fable 5 returned to general access on July 1. Mythos 5 has been reintroduced, but only to vetted U.S. organizations as part of an ongoing government review process, TechCrunch reported. It remains restricted for everyone else as of today.

Anthropic had marketed Mythos as a uniquely powerful and dangerous tool, something close to a "doomsday cybermachine" in TechCrunch's characterization, that could only be handed to carefully screened users under tight controls. That marketing may have made the export control response almost inevitable once a bypass was reported, whether or not the underlying jailbreak claim holds up.

Mark Dowd, a well-known security researcher who has spent decades finding and selling zero-day vulnerabilities to Western governments rather than reporting them to software makers, put the objection bluntly on a cybersecurity podcast, according to TechCrunch: "it's not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what's not."

Dowd's business model deserves context. Governments pay a premium for zero-days specifically because they stay unpatched, which makes them useful for intelligence and offensive operations rather than defense. Dowd acknowledged to TechCrunch that this work may bias his view of guardrails. But he's not alone in the complaint, and the researchers making it aren't all selling exploits to spy agencies. Some are consultants like Anley, whose entire job is finding bugs so they get fixed, not exploited.

There's a genuine tension AI companies haven't resolved. A model that will happily explain how to exploit a buffer overflow is useful to a penetration tester hired by a bank and useful to a criminal targeting that same bank. There's no clean technical way for the model to know which one is asking. Vetting programs are the current answer, but they add friction, delay, and a gatekeeping layer that didn't exist before AI models became part of the workflow.

The counterargument is straightforward: these are new, powerful tools, and some caution while the industry figures out how to screen users is reasonable given how fast AI capabilities have moved. Nobody disputes that a model with zero guardrails would be trivially useful to actual criminals. The disagreement is over where the line sits and who draws it.

What's unresolved is whether OpenAI and Anthropic will loosen default guardrails for a broader set of professional users, or whether vetting programs become the permanent price of entry for anyone doing serious offensive security work with AI assistance. Mythos 5's restricted status, limited to vetted U.S. organizations under an active government review as of this writing, suggests the government isn't in a hurry to answer that question either.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center-left
TechCrunchHow AI guardrails are impeding the work of offensive cybersecurity researchers