MM-LLMs: Recent Advances in MultiModal Large Language Models
Duzhen Zhang, Yahan Yu, Jiahua Dong, Chenxing Li, Dan Su, Chenhui Chu, Dong Yu
Tencent (China) Kyoto University Mohamed bin Zayed University of Artificial Intelligence
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
In the past year, MultiModal Large Language Models (MM-LLMs) have undergone substantial advancements, augmenting off-the-shelf LLMs to support MM inputs or outputs via cost-effective training strategies.The resulting models not only preserve the inherent reasoning and decision-making capabilities of LLMs but also empower a diverse range of MM tasks.In this paper, we provide a comprehensive survey aimed at facilitating further research on MM-LLMs.Initially, we outline general design formulations for model architecture and training pipeline.Subsequently, we introduce a taxonomy encompassing 126 MM-LLMs, each characterized by its specific formulations.Furthermore, we review the performance of selected MM-LLMs on mainstream benchmarks and summarize key training recipes to enhance the potency of MM-LLMs.Finally, we explore promising directions for MM-LLMs while concurrently maintaining a real-time tracking website 1 for the latest developments in the field.We hope that this survey contributes to the ongoing advancement of the MM-LLMs domain.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AINatural Language Processing Techniques
Topic Modeling
参考文献 0
引用本文 223
按被引量排序,此处列出前 3 条