When an agent does something it shouldn't, the trace tells you what it saw. It doesn't tell you which part of that made it act. The usual move is to read through the messages, pick the line that looks guilty, add a rule to the system prompt and rerun once. But one rerun can't tell you whether the rule worked or the model just went the other way... Weiterlesen
Intelligence View
⚡ tsecurity.de Intelligence
Finding the sentence that made an AI agent misbehave
When an agent does something it shouldn't, the trace tells you what it saw. It doesn't tell you which part of that made it act. The usual move is to read…