MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Bin Lin, Zhenyu Tang, Yang Ye, Jinfa Huang, Junwu Zhang, Yatian Pang, Peng Jin, Munan Ning 等 10 位
Peking University University of Rochester
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Recently, remarkable progress has been made in scaling up Large Language Models (LLMs) through the use of the sparse Mixture-of-Expert (MoE) layers without significantly increasing computational cost. However, the transition from a pre-trained LLM to a sparse Large Vision-Language Model (LVLM) with MoE remains an open challenge. Directly fine-tuning an LLM to a sparse LVLM often leads to training collapse, characterized by (1) a large modality feature distribution gap and (2) expert load imbalance. This paper proposes a three-stage decoupled weight training process. In the first two stages, the model learns to adapt the LLM to an LVLM. In the third stage, the FFN weights from the second stage are used as lossless initialization for expert weights, effectively constructing a sparse model with a vast number of parameters while maintaining constant computational cost. Through extensive ablation experiments, we derive three empirical guidelines and propose a sparse LVLM termedMoE-LLaVA. MoE-LLaVA is a MoE-based sparse LVLM architecture, which uniquely activates only the top-$k$experts through routers during deployment, keeping the remaining experts inactive. Extensive experiments demonstrate that MoE-LLaVA outperforms LLaVA-1.5-7B with an average improvement of 4.6 across nine visual understanding benchmarks. Notably, with only 2.2B active parameters, our MoE-LLaVA shows comparable result with LLaVA-1.5-13B (87.0 vs. 85.9) on POPE benchmark. Our work establishes a baseline for sparse LVLMs and provides empirical guidelines for exploring the sparse LVLMs. Our code is available at:https://github.com/PKU-YuanGroup/MoE-LLaVA.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIMultimodal Machine Learning Applications
Domain Adaptation and Few-Shot Learning · Advanced Neural Network Applications
参考文献 41
此处列出前 3 条
引用本文 32
按被引量排序,此处列出前 3 条