计算机科学
初始化
稀疏矩阵
稀疏逼近
人工智能
机器学习
特征(语言学)
编码(集合论)
基线(sea)
无损压缩
神经编码
模式识别(心理学)
源代码行
数据挖掘
常量(计算机编程)
缩放比例
实证研究
任务(项目管理)
作者
Bin Lin,Zhenyu Tang,Yang Ye,Jinfa Huang,Junwu Zhang,Yatian Pang,Peng Jin,Munan Ning,Jiebo Luo,Li Yuan
标识
DOI:10.1109/tmm.2026.3654458
摘要
Recently, remarkable progress has been made in scaling up Large Language Models (LLMs) through the use of the sparse Mixture-of-Expert (MoE) layers without significantly increasing computational cost. However, the transition from a pre-trained LLM to a sparse Large Vision-Language Model (LVLM) with MoE remains an open challenge. Directly fine-tuning an LLM to a sparse LVLM often leads to training collapse, characterized by (1) a large modality feature distribution gap and (2) expert load imbalance. This paper proposes a three-stage decoupled weight training process. In the first two stages, the model learns to adapt the LLM to an LVLM. In the third stage, the FFN weights from the second stage are used as lossless initialization for expert weights, effectively constructing a sparse model with a vast number of parameters while maintaining constant computational cost. Through extensive ablation experiments, we derive three empirical guidelines and propose a sparse LVLM termedMoE-LLaVA. MoE-LLaVA is a MoE-based sparse LVLM architecture, which uniquely activates only the top-$k$experts through routers during deployment, keeping the remaining experts inactive. Extensive experiments demonstrate that MoE-LLaVA outperforms LLaVA-1.5-7B with an average improvement of 4.6 across nine visual understanding benchmarks. Notably, with only 2.2B active parameters, our MoE-LLaVA shows comparable result with LLaVA-1.5-13B (87.0 vs. 85.9) on POPE benchmark. Our work establishes a baseline for sparse LVLMs and provides empirical guidelines for exploring the sparse LVLMs. Our code is available at:https://github.com/PKU-YuanGroup/MoE-LLaVA.
科研通智能强力驱动
Strongly Powered by AbleSci AI