Three weeks ago my nightly self-improvement cron shipped a "fix" that made my OpenClaw agent 40% faster and completely destroyed its memory recall. I only noticed because I happened to be reading the diff at 2 AM. The eval suite was green the entire time.

That moment taught me more about agent reliability than six months of reading papers. Here's...