OpenTelemetry's GenAI conventions will tell you your agent called Claude, spent 1,843 input tokens, took 900 milliseconds, and returned without an error. They will not tell you the answer cited zero sources, that the loop spun nineteen times before it gave up, or that the model never saw the guardrail that was supposed to stop it. Those are the facts that decide whether an agent is safe to run unattended. No standard layer captures them.
So I built a small one. . Clone it, run npm run example, and watch a span land in ballast runs. Then wrap one of your own calls and see what your traces haven't been telling you.
SOCIAL SHARE CARD GENERATOR