Signal/Power Integrity Aware Design and Computing Performance Analysis of SRAM-Bridge Embedded-GPU-HBM Architecture
Haeseok Suh, Jiwon Yoon, Hyunjun An, Keunwoo Kim, Taein Shin, Keeyoung Son, Seonguk Choi, Junghyun Lee 等 12 位
Korea Advanced Institute of Science and Technology Sejong University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
In this article, we propose the static RAM (SRAM)-bridge embedded-GPU-high bandwidth memory (SBE-GPU-HBM) architecture and analyze its signal integrity (SI), power integrity (PI), and computing performance. SBE-GPU-HBM is a next generation artificial intelligence (AI) accelerator architecture that embeds the SRAM-bridge (SB) chip within the interposer to reduce HBM data traffic and shorten the off-chip interconnect length. To maximize the memory bandwidth for the proposed architecture, the local silicon interconnect (LSI) stack-up is designed to reduce signal loss and crosstalk. Also, power distribution network (PDN) comprising power/ground planes and vias is designed to mitigate simultaneous switching noise (SSN). To validate the proposed design, LSI and PDN are co-analyzed in time domain to obtain the highest achievable bandwidth. Furthermore, the reduced HBM traffic by SRAM is quantified by using an analytical model-based simulator. By incorporating the enhanced bandwidth and reduced HBM traffic, the system-level computing performance is evaluated. The results demonstrate that the SBE-GPU-HBM can increase the memory bandwidth by up to 187% and reduce the HBM traffic by 17.1% during the large language model (LLM) inference decode process. Ultimately, the proposed architecture achieves a 216% improvement in computing performance for LLM inference relative to the conventional GPU-HBM architecture.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程3D IC and TSV technologies
Parallel Computing and Optimization Techniques · Interconnection Networks and Systems