Why AI Guardrails Are Crippling the Defenders Built to Protect Us.
When AI Guardrails Block the Good Guys,Too
Inside the vetting programs, export controls, and workarounds shaping how cybersecurity researchers actually use frontier AI.
2: Major vetted-access cyber programs: OpenAI & Anthropic
July 1: Date Fable 5 access was restored after export controls
0: Guardrails on the open-source models researchers fall back to
1: The Guardrails Built to Stop Hackers Are Also Stopping Defenders:
AI giants built strict guardrails and vetted-access programs to keep malicious hackers from weaponizing their models — but those same limits are now getting in the way of the legitimate researchers who defend systems for a living.
Offensive cybersecurity researchers proactively probe software for unknown vulnerabilities — "zero-days" — before criminals find them. Increasingly, they say frontier AI models refuse to help with core parts of that work. Mark Dowd, a veteran researcher who sells zero-days to Western governments, put it bluntly on a recent cybersecurity podcast: he's uncomfortable with large AI companies making unilateral calls about what counts as safe in security.
Chris Anley, chief scientist at security consulting firm NCC Group, said asking a model to try to exploit a bug is a key step in confirming a vulnerability is real and worth fixing — but when a guardrail makes the model refuse outright, it hurts defenders, not just attackers. As he described it, the same prompt that helps fix code is also a roadmap for finding critical vulnerabilities, and the two uses can't be cleanly separated.
2: Anthropic's Mythos, Export Controls, and the Vetting Bottleneck:
Anthropic's own marketing has amplified the tension.
In June, the U.S. government placed export control restrictions on Anthropic's Mythos and Fable models, a move at least partly prompted by a report claiming their guardrails could be bypassed to build and execute cyberattacks. Anthropic had promoted Mythos as a kind of doomsday cybermachine reserved for carefully vetted users under strict guardrails — a framing that raised the stakes when the export controls hit.
Those restrictions have since been lifted: Fable 5 returned to general access on July 1, while Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government's ongoing review.
That gatekeeping isn't unique to Mythos. Anthropic's Cyber Verification Program and OpenAI's Trusted Access for Cyber program both let researchers apply for access to versions of frontier models with fewer restrictions — but approval isn't guaranteed, and the guardrails inside those programs aren't always consistent.
I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program. Instead of analyzing a vulnerability and reasoning through the exploitability, you're trying to find why you're getting inconsistent results or why are models over-sanitizing the output. — Chris Thompson, CEO, RemoteThreat
3: Why Researchers Are Falling Back on Open-Source Models:

The Hidden AI War
Nobody Is Telling You About
Our latest documentary deep-dive into the geopolitical struggle for machine intelligence dominance. Explore the two paths of AI development: open source vs. closed architecture.
Faced with unpredictable refusals, many researchers route around the frontier labs entirely.
● Paolo Stagno of Crowdfense said his team uses frontier models only for reverse engineering, avoiding cloud-based AI for vulnerability discovery over concerns that sensitive data could leak or be absorbed into future training runs — for that work, they use open-source models run locally.

Move Over Nvidia: Anthropic and Samsung Team Up to Break the AI Chip Monopoly
● Giuseppe Cali, a security researcher who develops exploits, said guardrails don't affect him because he uses AI only for reverse engineering and tooling, not for the actual bug discovery he prefers to own himself.
Support our research
Independent analysis fueled by you.
● A researcher at a smartphone-component manufacturer not enrolled in Anthropic's CVP said tools become unusable the moment they detect security-related work.
● Chris Thompson said researchers are increasingly pushed toward Chinese open-source models like GLM, which can be downloaded and run locally with no vetting or usage restrictions at all.
Thompson called that shift concerning on its own terms: responsible researchers being pushed away from U.S.-governed systems toward foreign-owned ones. His view is that frontier labs should open up their programs and hold bad actors accountable rather than tightening restrictions further — otherwise, he argued, defenders risk losing the AI race just as attacks scale up in speed and volume.
4: The Real Lesson for Enterprise AI Adoption:
This isn't just a cybersecurity story — it's a preview of what happens when general-purpose AI guardrails collide with real-world, specialized work.
Frontier models are tuned to guard against the worst-case use of the entire internet, which means legitimate, narrow use cases inside a business can get caught in the same net as malicious ones. Enterprises don't need a model that second-guesses every request meant for the internet at large — they need AI that understands the specific, legitimate work it was built to do.
Enterprise AI Shouldn't Feel Like Negotiating With a Black Box.
The frontier labs' one-size-fits-all guardrails were built to police the entire internet's worth of use cases — which is exactly why security researchers describe spending more time negotiating with the model than doing their actual jobs.
Business AI shouldn't work that way. Otherworlds AI's Agent+ Business AI Platform is scoped to your workflows from the start, so it doesn't second-guess legitimate work the way general-purpose frontier models do, and custom enterprise AI builds let you define exactly what your AI is trusted to do.

Meta's Next Big Bet: This New App Lets You Build Games Simply by Typing a Prompt
Predictable behavior, built for your business, not a vetting queue.
Explore Agent+ at otherworldsai.com







