Anthropic has published new technical details on the cybersecurity safeguards protecting Claude Fable 5, which is now globally available following its redeployment. The disclosure covers two major initiatives: a breakdown of the model’s safety classifiers, and an early-draft framework for scoring the severity of AI jailbreaks developed in partnership with Anthropic’s Glasswing collaborators. Claude Fable […]
The post Anthropic Details Claude Fable 5 Cyber Safeguards and AI Jailbreak Severity Framework appeared first on Cyber Security News.