研究论文
HALO: A heterogeneous accelerator for low-latency and energy-efficient edge LLM inference
Kunming Zhang, Zhihua Fan, Yanhuan Liu, Lexin Wang, Yuqun Liu, Haibin Wu, Xiaochun Ye, Wenming Li
Institute of Computing Technology University of Chinese Academy of Sciences
来源Future Generation Computer Systems
年份2026
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
摘要 · 节选
暂未获取摘要。可打开原文或 PDF,后续可基于全文生成更完整的速读。
逐年被引趋势
暂无年度引用数据
关键指标
0
被引次数 · OpenAlex
0.00
领域内被引倍数
同类平均 = 1
同类平均 = 1
前 45%
引用位次
同领域 · 同年份 · 同类型
同领域 · 同年份 · 同类型
22
参考文献
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:文献信息
论文问答
当前基于文献信息回答
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIAdvanced Neural Network Applications
Adversarial Robustness in Machine Learning · Machine Learning and Data Classification
参考文献 22
Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable Architecture
被引 216Liqiang Lu, Yicheng Jin, Hangrui Bi · 2021
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
被引 113Thierry Tambe, Coleman Hooper, Lillian Pentecost · 2021
DOTA: detect and omit weak attentions for scalable transformer acceleration
被引 130Zheng Qu, Liu Liu, Fengbin Tu · 2022
此处列出前 3 条
引用本文 -
暂无引用本文的记录