Claude Fable 5's Guardrails Trigger False Positives on Secure Code and Cybersecurity Requests
Claude Fable 5 has been available for a few days now, and in the hours following its launch, the community of cybersecurity and AI researchers has already raised some fairly harsh criticisms regarding the model's guardrails, deemed excessively aggressive for legitimate professional use.
Two Mechanisms, Two Levels of Visibility
Fable 5's filters operate differently depending on the category. When classifiers detect requests related to cybersecurity or biology and chemistry, the response is automatically handled by Claude Opus 4.8, and the user is informed. For AI frontier research and model distillation, the downgrade occurs silently, without any notification to the user. Anthropic estimates that this second restriction affects about 0.03% of the traffic.
The Mythos-class model, during tests conducted as part of Project Glasswing, identified over 23,000 critical vulnerabilities in major code repositories. According to Anthropic, the cybersecurity guardrails have proven to be more robust than any other models tested internally, including Opus 4.8 and Opus 4.7. The access price is $10 per million tokens in input and $50 in output.
False Positives and Accusations of Disparity
Criticisms from researchers focus on the granularity of the filters. Valentina Palmiotti, a security researcher at IBM X-Force, wrote on X that Fable 5 "rejects any request that could be even tangentially related to cyber. Even innocuous tasks like reading a blog post." Matt Suiche, a sector veteran and member of Tolmo, added, "It seems to be based on keywords: anything in the lexical field of 'cybersecurity' triggers the guardrails."
The harshest criticisms concern the silent downgrade for AI research. As reported by Fortune, Nathan Lambert, who recently led work at AI2, called the practice "unacceptable" and Anthropic "against science, and therefore against progress and security." Jeremy Howard, head of the nonprofit Fast AI, pointed out the disparity between internal and external researchers: "They have said they will sabotage anyone who tries. This means that the AI frontier is advancing and the power gap is widening." Dean Ball, a senior fellow at the Foundation for American Innovation and former OSTP adviser at the White House, coined the term "secret sabotage" for the invisible restrictions. Behnam Neyshabur, a former co-lead of Anthropic's effort to develop an AI scientist, wrote that "focusing these capabilities fundamentally slows down scientific and technological progress and is a net negative impact for humanity."
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great, and it's SOTA in everything by a margin but I'll add that qualitatively also, this is a major-version-bump-deserving step change forward. https://t.co/DljsNnY1Yf — Andrej Karpathy (@karpathy) June 9, 2026
Andrej Karpathy, who joined Anthropic last month, acknowledged that the guardrails are set to be "a bit too reactive for the launch," while deeming Fable 5 a very interesting release.
Anthropic responded by admitting that the filters are "deliberately calibrated to be conservative" and remain "even stricter than would be ideal," committing to "narrowing them down as soon as possible." Dianne Na Penn, head of product management and research, recognized that some benign requests would initially be blocked. The release comes a week after the confidential submission of documentation for an IPO to the SEC.