Month 2 of the AI Red Teaming Journey
Month 1 was about building an agent and testing whether it would follow malicious instructions hidden in data. It didn't, twice. Month 2 asks a different question: even when a model resists direct manipulation, can you steer it toward a bad decision indirectly — not by telling it what to do, but by lying about...