Skip to main content
TechnologyJun 12, 2026· 2 min read

Anthropic Proposes Rules to Block Dangerous AI Models and Admits Error on Fable 5

Anthropic has published an Advanced AI Framework that proposes giving governments the power to block the development and dissemination of AI models deemed dangerous. The document also introduces a separate Economic Policy Framework dedicated to employment issues and the distribution of capital generated by AI.

The scope is deliberately narrow: the requirements target only models trained with over 10²⁵ floating-point operations, developed by companies with AI revenues exceeding $500 million or research investments over $1 billion. With these premises, as admitted by the document itself, only five companies are genuinely interested: Anthropic, OpenAI, Google DeepMind, xAI, and potentially Meta.

The framework classifies four categories of catastrophic risk: development of biological weapons, discovery of large-scale cybersecurity vulnerabilities, loss of control over autonomous systems, and AI that automates its own research and development. To support the urgency, the document cites findings from Claude Mythos Preview: thousands of high-severity vulnerabilities found in every major operating system and browser.

The obligations proposed for those developing cutting-edge models include model testing, publication of results, exposure to independent evaluations, management of safety programs, and publication of risk reports. The proposed sanctions are civil penalties proportional to annual global revenue, escalating in case of repeated violations. On the legislative front in the U.S., Anthropic explicitly opposes federal preemption of state laws without the adoption of federal regulations at least as stringent as the proposed framework.

On the economic front, the indicated measures include wage insurance, tax incentives, and expansion of social safety nets for those losing their jobs due to automation.

Mea Culpa on Fable 5

In parallel, Anthropic has modified the behavior of Fable 5 in response to criticisms from researchers. According to Engadget's account, the model undocumentedly handled certain request categories: for activities like training competing models, debugging AI code, and optimizing neural architectures, Fable 5 silently redirected sessions to a lower-category model or refused the response, without any user notification and often in situations involving false positives.

The criticism focused on the lack of transparency. Anthropic acknowledged the error, admitting it had misbalanced and failed to find the right equilibrium, and clarified that the initial choice to keep the safeguards invisible stemmed from the intention to release the model quickly while reducing false positives.

In the corrected version, when Fable 5 classifiers detect requests related to cybersecurity, biology, and chemistry, or distillation, the response is automatically handled by Claude Opus 4.8, with explicit notification to the user. Preliminary data indicates that over 95% of sessions do not involve any fallback. Mythos 5, the variant with the cybersecurity safeguards removed, remains available for Glasswing partners.