Jenga: Effective Memory Management for Serving LLM with Heterogeneity
Chen Zhang, Kuntai Du, Shu Liu, Woosuk Kwon, Xiangxi Mo, Yufeng Wang, Xiaoxuan Liu, Kaichao You 等 13 位
University of California, Berkeley Tsinghua University University of Chicago College of San Mateo
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Large language models are widely used but expensive to run. To reduce costs, it is crucial to maximize request batch size through efficient GPU memory management. Existing approaches, such as PagedAttention, struggle with modern LLMs because of the growing heterogeneity in the sizes of models' internal embeddings and attention mechanisms.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
物理Oil Palm Production and Sustainability
Lipid metabolism and biosynthesis · Coconut Research and Applications
参考文献 0
引用本文 3
按被引量排序,此处列出前 3 条