How AI guardrails are impeding the work of offensive cybersecurity researchers
For several months, leading AI companies have implemented carefully vetted programs and stringent safeguards to prevent their models from being exploited by malicious hackers. However, these restrictions are increasingly obstructing the efforts of legitimate network defenders and offensive cybersecurity researchers. In June, the U.S. government imposed export control restrictions on Anthropic’s AI models Mythos and Fable, partially triggered by a report suggesting that the models’ guardrails designed to block malicious cyberattacks could be circumvented. Although the export controls on Fable 5 and Mythos 5 have since been lifted—with Fable 5 resuming general access and Mythos 5 made available only to vetted U.S. organizations under review—the perception of these models as potentially dangerous “doomsday” machines continues to influence access policies.
This type of controlled access is not unique to Mythos. Both Anthropic and OpenAI offer cybersecurity researchers special programs—Anthropic’s Cyber Verification Program and OpenAI’s Trusted Access for Cyber—that allow approved users to access models with reduced restrictions. Yet, these guardrails have drawn criticism, especially from researchers whose roles involve identifying unknown system vulnerabilities and testing exploit possibilities prior to criminal discovery. Mark Dowd, a noted security researcher, expressed discomfort at the notion that large corporations are unilaterally determining what constitutes secure or unsafe use of these technologies. Dowd, who has a history of finding and selling zero-day exploits to governments, acknowledged potential bias but is representative of many in offensive cybersecurity who rely on AI tools and face guardrail challenges.
Chris Anley from NCC Group highlighted that using AI to attempt bug exploitation is crucial for verifying legitimate vulnerabilities that merit remediation. However, when guardrails prevent AI from responding to certain queries, they impede defenders’ ability to identify critical flaws. Anley likened the AI tool to a hammer—essential for construction but also inherently capable as a weapon—illustrating how the same technology serves both offensive and defensive purposes that cannot easily be separated. When encountering obstacles, some researchers resort to open source AI models without guardrails. Similarly, Paolo Stagno of Crowdfense criticized AI vendors for treating users like children needing supervision through stringent vetted programs, noting his preference to use frontier AI models only for reverse engineering rather than vulnerability detection or exploit creation due to concerns about data privacy when connected to cloud-based models.
On the other hand, some researchers, such as Giuseppe Cali, are not hindered by guardrails because they use AI mainly for code analysis and tool creation rather than offensive work. Cali appreciates the speed AI provides for preliminary tasks while maintaining personal control over vulnerability discovery and weaponization, underscoring a desire to preserve the craft of bug hunting. However, others report less success; an anonymous researcher from a smartphone component company disclosed that without participation in Anthropic’s Cyber Verification Program, their tools are essentially ineffective for security work due to overly restrictive guardrails. Chris Thompson, CEO of RemoteThreat, noted that guardrails can act inconsistently and change daily, frustrating researchers who find themselves negotiating AI limitations rather than focusing on vulnerability assessment. He pointed out that this drives experts toward using foreign open source models with no restrictions, such as China’s GLM, which raises concerns about pushing responsible researchers away from U.S.-governed systems.
Thompson advocates for increased transparency and broader access through AI programs rather than tightening constraints, coupled with accountability for misuse. He warns that without such measures, defenders risk falling behind as a wave of rapid and large-scale cyberattacks approaches. In his view, current restrictive policies stifle the very security professionals striving to counter emerging threats, undermining efforts to maintain cybersecurity in the face of escalating AI-driven attacks.