Key Takeaways
- Anthropic released Claude Fable 5, a Mythos‑class AI model with built‑in safeguards that block its use in high‑risk domains such as cybersecurity and biology.
- When a request touches a restricted area, Fable 5 automatically falls back to the less capable Claude Opus 4.8 to prevent misuse.
- Early data show that ≥95 % of sessions run entirely on Fable 5 without triggering a fallback, indicating the safeguards work in practice.
- Safety was validated through extensive internal red‑teaming, an external bug‑bounty program (>1,000 hours), and independent external red‑teaming—none uncovered universal jailbreaks.
- Trusted cybersecurity partners in Project Glasswing are being upgraded from Claude Mythos Preview to the full Claude Mythos 5, with plans to expand the program to ~150 new organizations.
- Both Fable 5 and Mythos 5 are priced at $10 per million input tokens and $50 per million output tokens, and Fable 5 is immediately available via the Claude API.
- Anthropic stresses that the uplift in capabilities could attract adversaries, motivating continual reinforcement of its safety barriers.
Introduction to Claude Fable 5
Anthropic unveiled Claude Fable 5 on Tuesday, marking the first time a Mythos‑class AI model of this capability level has been deemed safe enough for broad public and developer release. Positioned as a powerful general‑purpose model, Fable 5 outperforms its predecessors in software engineering, knowledge work, vision processing, and long‑running tasks. Yet, the company’s primary focus during development was safety, leading to the implementation of targeted usage restrictions that automatically curb the model’s output in domains deemed high‑risk.
Safety‑First Design: Domain‑Specific Blocks
To mitigate potential harm, Anthropic engineered Fable 5 to recognize and avoid generating content that could facilitate cyberattacks, biological weapon development, or other dangerous activities. When a user query touches one of these sensitive areas, the model does not simply refuse; instead, it falls back to Claude Opus 4.8, a less capable but safer predecessor. This fallback mechanism ensures that even if a request attempts to probe restricted knowledge, the system defaults to a model with insufficient capacity to produce actionable harmful outputs.
Empirical Evidence of Effectiveness
Early usage telemetry indicates that at least 95 % of sessions operate entirely on Fable 5’s full capabilities without invoking the fallback. This high success rate suggests that the safeguards are not overly restrictive for legitimate use cases while still catching the majority of attempts to venture into prohibited territory. Anthropic notes that the remaining 5 % of sessions likely involve edge‑case queries that the classifiers flag as potentially risky, triggering the safer Opus 4.8 pathway.
Adversarial Motivation and Continuous Vigilance
Anthropic acknowledges that the uplift from Mythos‑level capabilities is attractive to adversaries, especially those seeking financial gain from cyberattacks. Consequently, the company expects motivated actors to attempt circumventing its safety measures. This anticipation drives a continual cycle of testing, monitoring, and upgrading the safeguards to stay ahead of emerging jailbreak techniques.
Internal Red‑Teaming Rigor
Before release, Anthropic conducted extensive internal red‑teaming of the classifiers that govern the fallback logic. Engineers and security specialists attempted to craft prompts that would bypass the domain‑specific blocks, probing for weaknesses in the model’s decision‑making boundaries. The internal effort produced no universal jailbreak, reinforcing confidence in the robustness of the initial safety layer.
External Bug‑Bounty Program
Following internal validation, Anthropic launched an external bug‑bounty initiative that attracted security researchers worldwide. Over 1,000 hours of focused effort were devoted to discovering any method to coax Fable 5 into generating restricted outputs. The program concluded with zero universal jailbreaks reported, indicating that the external community also failed to find a reliable bypass.
Independent External Red‑Teaming Confirmation
To further substantiate the claims, Anthropic enlisted independent external red‑teams to perform adversarial testing without prior knowledge of the internal safeguards. These teams likewise uncovered no critical bypasses, underscoring the model’s resistance to coordinated attempts to elicit dangerous behavior. The convergence of internal, bounty‑driven, and independent testing results provides a layered validation of the safety architecture.
Project Glasswing: Privileged Access for Trusted Partners
Parallel to the public rollout, Anthropic announced that its Project Glasswing collaborators—trusted cybersecurity organizations—are being upgraded from the earlier Claude Mythos Preview to the full Claude Mythos 5 model. This upgrade grants these partners higher‑privilege access to the model’s capabilities while still operating under Anthropic’s overarching safety framework. The company plans to gradually expand this trusted‑access program through a structured vetting process, ensuring that only vetted entities receive the elevated privileges.
Expansion of Project Glasswing
Anthropic revealed that it is adding roughly 150 new organizations to Project Glasswing. While the complete list remains undisclosed, several notable cybersecurity and tech firms have already publicized their participation, including Dragos, Tenable, TrendAI (Trend Micro), Netskope, BeyondTrust, Rubrik, BT, Intercontinental Exchange, and Hitachi. Their involvement signals strong industry confidence in Anthropic’s safety‑first approach and the utility of Mythos 5 for advanced threat‑analysis, vulnerability research, and defensive tooling.
Pricing Model and Availability
Both Claude Fable 5 and the upgraded Mythos 5 are offered under a token‑based pricing scheme: $10 per million input tokens and $50 per million output tokens. Fable 5 is immediately accessible via the Claude API for developers, allowing seamless integration into existing workflows. The pricing reflects the model’s high capability tier while remaining competitive within the premium AI market segment.
Related Developments and Industry Impact
Anthropic’s announcement coincides with several related narratives in the AI security space. Reports highlight how Claude Mythos can compress N‑day exploit timelines into N‑hours, underscoring the dual‑use nature of powerful language models. Other discussions explore cryptographic invisibility techniques to shield AI‑built applications from reverse engineering, and debate whether AI will disrupt traditional bug‑bounty economies by automating vulnerability discovery. These conversations illustrate the broader tension between advancing AI capabilities and mitigating their potential for misuse.
Conclusion
Claude Fable 5 represents a significant milestone: a frontier‑class AI model that couples state‑of‑the‑art performance with rigorously engineered safeguards. By automatically deferring to a safer fallback in high‑risk contexts, subjected to exhaustive internal and external testing, and offering privileged access through a vetted partner program, Anthropic seeks to balance innovation with responsibility. As the model becomes widely available via the Claude API, its real‑world impact will hinge on continued vigilance, transparent reporting of safety metrics, and the collaborative effort of the AI community to uphold ethical boundaries.

