TranCIM: Full-Digital Bitline-Transpose CIM-based Sparse Transformer Accelerator With Pipeline/Parallel Reconfigurable Modes
Fengbin Tu, Zihan Wu, Yiqi Wang, Ling Liang, Liu Liu, Yufei Ding, Leibo Liu, Shaojun Wei 等 10 位
Tsinghua University University of California, Santa Barbara
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Transformer models achieve excellent results in the fields like natural language processing, computer vision, and bioinformatics. Their large numbers of matrix multiplications (MMs) lead to substantial data movement and computation. Although computing-in-memory (CIM) has proven to be an efficient architecture for MM computation, transformer’s attention mechanism raises new challenges in memory access and computation aspects: the dynamic MM in attention layers causes redundant OFF-chip memory access; Attention layers dominate transformer’s computation and require high precision. Thus, we design a bitline-transpose CIM-based transformer accelerator TranCIM with pipeline/parallel reconfigurable modes. The pipeline mode alleviates off-chip access for attention layers. The parallel mode is used by fully-connected (FC) layers for high parallelism. The full-digital CIM supports INT16 for attention layers and INT8 for FC layers, without analog CIM’s nonideal issues. Moreover, a sparse attention scheduler (SAS) is proposed to reduce attention computation. The fabricated TranCIM chip only consumes 15.59$\mu \text{J}$/Token for the bidirectional encoder representations from transformer (BERT)-base model, achieving$12.08\times $–$36.82\times $lower energy than prior CIM-based accelerators.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程Analog and Mixed-Signal Circuit Design
Photonic and Optical Devices · Advancements in PLL and VCO Technologies
参考文献 32
此处列出前 3 条
引用本文 93
按被引量排序,此处列出前 3 条