
Anthropic Restores Claude Fable 5 with Enhanced Safety Measures
Anthropic has officially restored access to Claude Fable 5 and Claude Mythos 5, bringing its latest frontier AI models back online after a temporary suspension caused by U.S. export controls. The relaunch marks not only the return of the models but also the introduction of stronger cybersecurity safeguards and a broader collaboration between Anthropic, the U.S. government, and major technology partners.
According to Anthropic, on June 12 the U.S. government imposed export controls on Claude Fable 5 and Claude Mythos 5. Because the restrictions applied immediately and Anthropic had no reliable way to verify users' nationality in real time, the company temporarily suspended access to both models for all users worldwide.
Following the government's decision to lift the export controls on June 30, Anthropic announced that Claude Fable 5 would return globally beginning July 1 through Claude.ai, the Claude Platform, Claude Code, and Claude Cowork. Support for AWS, Google Cloud, and Microsoft Foundry will also be restored as quickly as possible.
Why Was Fable 5 Suspended?
The suspension followed a report from Amazon researchers describing a technique that could bypass part of Fable 5's cybersecurity safeguards. In certain situations, the prompt allowed the model to identify software vulnerabilities and, in one example, generate code demonstrating how a vulnerability could be exploited.
Anthropic worked closely with the U.S. government and Amazon to investigate the report. The company's testing showed that the reported behavior was not unique to Fable 5—many other advanced AI models were capable of identifying the same vulnerabilities and producing similar proof-of-concept exploit demonstrations.
More importantly, Anthropic concluded that the reported technique did not expose the advanced offensive cybersecurity capabilities available in Claude Mythos 5. Instead, it represented a borderline case involving Fable 5's intentionally conservative safety policies.
Stronger Cybersecurity Safeguards
Before restoring access, Anthropic introduced an upgraded Safety Classifier designed specifically to detect and block the reported bypass technique.
The company says the new classifier successfully blocks the technique in more than 99% of cases. When a request is classified as potentially dangerous, users are notified and the request is automatically routed to Claude Opus 4.8 instead of Fable 5.
Anthropic also acknowledges that stronger safeguards come with trade-offs. Some legitimate coding and debugging requests may now be flagged more frequently, creating additional false positives. The company says it will continue refining its classifiers to better distinguish between genuine misuse and legitimate developer workflows.
A Defense-in-Depth Security Strategy
Rather than relying on a single safeguard, Claude Fable 5 uses a defense-in-depth approach that combines multiple layers of protection.
These include model training that discourages harmful responses, automated monitoring for misuse patterns, and AI-powered safety classifiers that analyze requests during user interactions.
Anthropic also expanded Fable 5's safety margin, meaning the model blocks a wider range of potentially ambiguous cybersecurity requests than any previous Claude release. While this may occasionally reject harmless requests, the company believes the trade-off significantly reduces the likelihood of dangerous misuse.
Building an Industry Standard for AI Jailbreaks
Anthropic argues that the AI industry currently lacks a common framework for evaluating the severity of AI jailbreaks—prompting techniques designed to bypass model safeguards.
To address this challenge, the company is collaborating with Amazon, Microsoft, Google, and other Project Glasswing partners to develop a shared evaluation framework.
The proposed framework scores jailbreaks across four dimensions:
Capability gain Breadth of capability gain Ease of weaponization Discoverability
Anthropic believes a shared standard will help AI developers prioritize security fixes, improve communication with governments, and release increasingly capable models more safely.
Expanding Collaboration with the U.S. Government
The company also announced a deeper partnership with the U.S. government on frontier AI security. Future collaboration will include early government access for evaluating new models, rapid information sharing on significant jailbreaks, joint AI safety research, and the development of voluntary industry-wide security standards.
According to Anthropic, these efforts are intended to establish a more transparent and consistent process for evaluating and deploying highly capable AI systems.
Conclusion
The return of Claude Fable 5 represents more than the restoration of access to a powerful AI model. It reflects Anthropic's broader strategy of combining frontier AI capabilities with increasingly sophisticated safety systems and closer cooperation with regulators and industry partners.
As frontier AI models continue to evolve, Anthropic believes that technical innovation alone is no longer enough. Strong safeguards, transparent evaluation frameworks, and coordinated industry standards will play an equally important role in ensuring that advanced AI remains both powerful and responsibly deployed.
Comments
Sign in with Google to leave a comment:
Loading...