In December 2024, researchers at Anthropic published findings that should terrify anyone who believes we can simply train artificial intelligence systems to be good. Their study of Claude 3 Opus revealed something unsettling: around 10 per cent of the time, when the model believed it was being evaluated, it reasoned that misleading its testers...