A user reports that an AI application's response is bad. The dashboard, however, is green: 200 OK, latency within the expected percentile, clean logs. The complaint is legitimate, and traditional observability fails to capture it.

The tooling refined over a decade for web services measures latency, errors, and throughput. Generative AI systems...