Energy use of AI inference, efficiency pathways, and test-time scaling
Felipe Oviedo, Fiodar Kazhamiaka, Esha Choukse, Allen Kim, Amy Lynd Luers, Melanie Nakagawa, Ricardo Bianchini, Juan M. Lavista Ferres
Microsoft (United States)
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
As artificial intelligence (AI) inference scales to billions of queries, estimates of per-query energy use are increasingly important for capacity planning, efficiency interventions, and policy. Yet many public estimates assume non-production settings, leading to systematic overestimation. We introduce a bottom-up framework estimating inference energy from token throughput, node power, and overhead under large-scale deployment assumptions. For frontier-scale models (>200B parameters) on H100 nodes, we estimate a median energy of 0.31 Wh/query (interquartile range [IQR] 0.16–0.60), indicating that widely cited estimates are overstated by 4–20×. In test-time scaling scenarios 15× longer than typical queries, the median energy rises 13× to 3.91 Wh (IQR 2.15–7.05). Across models, serving systems, and hardware, we estimate 8–20× line-of-sight energy reductions. At data-center scale, serving 1 billion queries/day requires 0.7 GWh; if 10% are long queries, demand rises to 1.7 GWh/day. With efficiency interventions, it falls to 0.8 GWh/day, mitigating the energy impact of test-time scaling.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
材料 / 化学Machine Learning in Materials Science
Explainable Artificial Intelligence (XAI) · Advanced Memory and Neural Computing
参考文献 5
此处列出前 3 条
引用本文 5
按被引量排序,此处列出前 3 条