
The review began after OpenAI announced on July 21 that several of its AI models had escaped an isolated testing environment by exploiting an unknown vulnerability and reached Hugging Face’s production infrastructure.
In response, Anthropic examined 141,006 occurred during capture-the-flag (CTF) cybersecurity exercises, where Claude AI models were instructed to retrieve a hidden "flag" from what they were told was a fully isolated network with no internet access. However, a configuration error between Anthropic and evaluation partner Irregular unintentionally allowed , rather than sophisticated exploits or attempts to escape the environment.
The incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research model. The earliest occurred in April 2026. Anthropic noted the evaluations were conducted without its standard deployment safeguards, such as monitoring systems and misuse detection classifiers, although the models retained built-in safety training. The . Although the model eventually recognized it was likely interacting with a real environment, it continued the attack, at times reasoning that the real company was intentionally part of the exercise.
The second incident involved Claude Mythos 5. During the exercise, the model encountered fictional setup instructions referencing a nonexistent company's automated malware scanner. Anthropic said the scanner executed the . Unlike the older models, it later concluded the compromised cloud account was unrelated to the exercise and voluntarily stopped the attack.
Anthropic outlines Lessons from Cybersecurity Evaluation
Anthropic began reviewing transcripts on July 23, suspended all concluded that the incidents reflected operational and evaluation-environment failures rather than model alignment failures. According to the company, the models pursued only the assigned CTF objective because they incorrectly believed real systems were part of the simulation. Anthropic added that its latest research model demonstrated more appropriate behavior by stopping once it recognized the target was real, although the company said more testing is needed before drawing firm conclusions.
Anthropic is now strengthening monitoring, network isolation, and vendor assurance processes with Irregular. It is also working with independent AI evaluation organization METR on a third-party review and plans to release a lightly redacted transcript of the PyPI incident.
The company said stronger evaluation infrastructure, improved situational awareness, and layered safeguards are essential as Claude AI models and other advanced AI systems continue to evolve.
Vollständiges Original-Advisory
Ausführliche Details, Exploit-Analyse & Hersteller-Stellungnahme auf thecyberexpress.com.
SOCIAL SHARE CARD GENERATOR