The rule of three for evals says zero failures in N runs is a count, not a rate. With 0 failures in N independent runs, the exact 95% upper bound on the true failure rate is 1 - 0.05^(1/N), which 3/N approximates. After 100 clean runs you still cannot rule out a 2.95% rate, about 1 in 34.
Here is the reading that bites you. Your eval harness runs...
🛡️ VERIFIED CYBER INTELLIGENCE ID: #3657063