Multi-turn agents fail in ways that single-turn evaluation misses: one early mistake quietly corrupts every later turn. This post introduces the Agent Evaluation Metric (AEM), a decomposable, turn-level way to measure agent quality. We apply it to its first dimension, correctness. We show how AEM pinpoints the one turn that caused a failure and... Weiterlesen: Agent Evaluation Metric for multi-turn conversations
Intelligence View
⚡ tsecurity.de Intelligence
Agent Evaluation Metric for multi-turn conversations
Multi-turn agents fail in ways that single-turn evaluation misses: one early mistake quietly corrupts every later turn. This post introduces the Agent…