The ability of AI models to perform end-to-end, multi-stage penetration tests that match the capabilities of humans undertaking the same tasks has improved dramatically in recent months, according to new benchmarks published by the UK government’s AI Security Institute (AISI).
In November 2025, the difficulty of cyber tasks the best models could complete was doubling every eight months, according to AISI, a research organization within the Department for Science, Innovation and Technology (DSIT).
By February this year, the performance improvements had accelerated, with the difficulty of the tasks AI models could complete doubling every 4.7 months, and since then the latest Claude Mythos Preview and GPT-5.5 models are showing even greater capability, warning businesses of the growing cyber security risks posed by AI models.
What’s clear is that the capabilities of AI models under real-world scenarios are rapidly improving and, on the evidence of the recent found the models could be error-prone and unreliable, especially for longer tasks.
of Claude Mythos that found mixed performance at some tasks. “How these known model limitations will actually limit real-world autonomous offensive campaigns is still being determined, but it does point to the need for a sophisticated validation harness to truly see the ceiling of model capabilities.”
According to Chris Lentricchia, director cloud and AI security strategy at Sweet Security, enterprises should also look at the upside — AI models aid attackers, but also defenders.
“This is not purely an offensive story. The same acceleration improving attacker capability can also improve defensive capability in areas like proactive threat detection and response automation. Benchmarks are best viewed as indicators for understanding whether enterprise defenses are evolving fast enough to keep pace with accelerating AI capability,” said Lentricchia.
SOCIAL SHARE CARD GENERATOR