How AI guardrails are impeding security researchers
AI labs built strict safety filters to stop malicious hacking, but those same limits now slow legitimate defenders. Researchers told TechCrunch how guardrails are impeding vulnerability checks and exploit analysis on OpenAI and Anthropic models, forcing many toward local open-source systems with fewer restrictions.
Key Takeaways
- OpenAI's Trusted Access for Cyber and Anthropic's Cyber Verification Program ease some limits for vetted users.
- U.S. export controls hit Anthropic's Mythos and Fable in June 2026; Fable 5 returned to general access on July 1.
- Security leaders say refusals block confirming bugs and burn time negotiating with models.
- Some researchers pivot to local open-source models, including Chinese systems like GLM.
- Others still use frontier AI only for reverse engineering, not exploit building.
Why are AI companies locking down cyber features?
For months, OpenAI and Anthropic have added guardrails and vetted programs to keep malicious hackers from abusing their models. In June 2026, the U.S. government also imposed export controls on Anthropic's Mythos and Fable models after reports that their cyber safety limits could be bypassed.
Those export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1, while Mythos 5 was reintroduced only to vetted U.S. organizations under government review, according to TechCrunch.
Anthropic has marketed Mythos as a powerful cyber-capable system that should reach carefully vetted users only. Both firms still offer moderated access paths: OpenAI's Trusted Access for Cyber and Anthropic's Cyber Verification Program.
How do researchers say the limits hurt their work?
Offensive cybersecurity researchers hunt unknown flaws and build exploits before criminals do. Several told TechCrunch the same refusal systems now get in their way.
Chris Anley, chief scientist at NCC Group, said asking a model to try exploiting a bug helps confirm a real vulnerability worth fixing. When a guardrail refuses that prompt, defenders lose a useful check. He compared AI to a hammer: essential for building, but also a weapon that cannot be fully split into "safe" and "unsafe" uses.
Mark Dowd, a veteran zero-day researcher, said he is uneasy that large companies make "arbitrary decisions about what is safe in security." Paolo Stagno of CrowdFense argued AI firms "essentially treat customers like children who need babysitting."
Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, said guardrails can be inconsistent day to day—even inside vetted programs. Researchers then spend time negotiating with the model instead of analyzing exploitability.
Are researchers abandoning U.S. frontier models?
When blocked, Anley's team sometimes falls back to open-source models with no guardrails. Stagno said his group uses frontier models for reverse engineering but keeps vulnerability and exploit work on local open-source systems to avoid leaking sensitive data into cloud training runs.
An anonymous researcher at a smartphone-component maker said his employer is not in Anthropic's CVP program, so tools stop when they detect security-related work. Thompson warned that responsible researchers are being pushed toward foreign-owned systems such as Chinese open-source GLM models that run locally with no vetting.
Not everyone feels blocked the same way. Giuseppe Cali said guardrails are not impeding his work because he uses AI for reverse engineering and tooling, not bug discovery or weaponization.
What do experts want instead of tighter locks?
Thompson urged frontier labs to open responsible access and hold abusers accountable, arguing defenders will otherwise lose the AI race as attacks accelerate. For more on AI's collision with security and emerging tech, browse our Future Tech & AI Wonders coverage.