研究论文
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Yanwei Li, Chengyao Wang, Jiaya Jia
TDK (China) University of Hong Kong
来源Lecture notes in computer science
年份2024
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
摘要 · 节选
暂未获取摘要。可打开原文或 PDF,后续可基于全文生成更完整的速读。
逐年被引趋势
79400
24
7925
26
关键指标
119
被引次数 · OpenAlex
28.16
领域内被引倍数
同类平均 = 1
同类平均 = 1
前 0.3%
引用位次
同领域 · 同年份 · 同类型
同领域 · 同年份 · 同类型
39
参考文献
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:文献信息
论文问答
当前基于文献信息回答
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIMultimodal Machine Learning Applications
Topic Modeling · Natural Language Processing Techniques
参考文献 39
ActivityNet: A large-scale video benchmark for human activity understanding
被引 2,671Fabian Caba Heilbron, Víctor Escorcia, Bernard Ghanem · 2015
ReferItGame: Referring to Objects in Photographs of Natural Scenes
被引 1,107Sahar Kazemzadeh, Vicente Ordóñez, Mark Matten · 2014
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
被引 5,286Ranjay Krishna, Yuke Zhu, Oliver Groth · International Journal of Computer Vision · 2017
此处列出前 3 条
引用本文 119
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
被引 82Liang Chen, Haozhe Zhao, Tianyu Liu · Lecture notes in computer science · 2024
VideoChat: chat-centric video understanding
被引 67Kunchang Li, Yinan He, Yi Wang · Science China Information Sciences · 2025
Uni-MoE: Scaling Unified Multimodal LLMs With Mixture of Experts
被引 49Yunxin Li, Shenyuan Jiang, Baotian Hu · IEEE Transactions on Pattern Analysis and Machine Intelligence · 2025
按被引量排序,此处列出前 3 条