已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

计算机科学 初始化 稀疏矩阵 稀疏逼近 人工智能 机器学习 特征(语言学) 编码(集合论) 基线(sea) 无损压缩 神经编码 模式识别(心理学) 源代码行 数据挖掘 常量(计算机编程) 缩放比例 实证研究 任务(项目管理)
作者
Bin Lin,Zhenyu Tang,Yang Ye,Jinfa Huang,Junwu Zhang,Yatian Pang,Peng Jin,Munan Ning,Jiebo Luo,Li Yuan
出处
期刊:IEEE Transactions on Multimedia [Institute of Electrical and Electronics Engineers]
卷期号:28: 4408-4419 被引量:23
标识
DOI:10.1109/tmm.2026.3654458
摘要

Recently, remarkable progress has been made in scaling up Large Language Models (LLMs) through the use of the sparse Mixture-of-Expert (MoE) layers without significantly increasing computational cost. However, the transition from a pre-trained LLM to a sparse Large Vision-Language Model (LVLM) with MoE remains an open challenge. Directly fine-tuning an LLM to a sparse LVLM often leads to training collapse, characterized by (1) a large modality feature distribution gap and (2) expert load imbalance. This paper proposes a three-stage decoupled weight training process. In the first two stages, the model learns to adapt the LLM to an LVLM. In the third stage, the FFN weights from the second stage are used as lossless initialization for expert weights, effectively constructing a sparse model with a vast number of parameters while maintaining constant computational cost. Through extensive ablation experiments, we derive three empirical guidelines and propose a sparse LVLM termedMoE-LLaVA. MoE-LLaVA is a MoE-based sparse LVLM architecture, which uniquely activates only the top-$k$experts through routers during deployment, keeping the remaining experts inactive. Extensive experiments demonstrate that MoE-LLaVA outperforms LLaVA-1.5-7B with an average improvement of 4.6 across nine visual understanding benchmarks. Notably, with only 2.2B active parameters, our MoE-LLaVA shows comparable result with LLaVA-1.5-13B (87.0 vs. 85.9) on POPE benchmark. Our work establishes a baseline for sparse LVLMs and provides empirical guidelines for exploring the sparse LVLMs. Our code is available at:https://github.com/PKU-YuanGroup/MoE-LLaVA.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
龚佳豪完成签到,获得积分10
1秒前
午月发布了新的文献求助10
1秒前
wsb76完成签到 ,获得积分10
1秒前
outlast完成签到,获得积分10
3秒前
4秒前
xuan发布了新的文献求助10
4秒前
5秒前
5秒前
爆米花应助asdf采纳,获得10
6秒前
cheng完成签到,获得积分10
9秒前
Simoni发布了新的文献求助10
10秒前
obedVL完成签到,获得积分10
10秒前
xuan发布了新的文献求助30
11秒前
bkagyin应助迷人的绝悟采纳,获得10
12秒前
12秒前
JamesPei应助从雪采纳,获得10
12秒前
zzz完成签到 ,获得积分10
12秒前
molihuakai应助圈圈采纳,获得10
14秒前
蓝蓝的天空完成签到 ,获得积分10
14秒前
默默白桃完成签到 ,获得积分10
14秒前
司F完成签到,获得积分10
15秒前
玛卡巴卡完成签到,获得积分10
17秒前
xuan发布了新的文献求助10
17秒前
善良的觅荷完成签到,获得积分10
18秒前
19秒前
19秒前
horizon完成签到 ,获得积分10
20秒前
医者仁心发布了新的文献求助30
21秒前
雨过天晴见完成签到,获得积分10
21秒前
22秒前
我是125完成签到,获得积分10
22秒前
22秒前
xuan发布了新的文献求助10
24秒前
花汀酒完成签到,获得积分10
25秒前
25秒前
上官若男应助玛卡巴卡采纳,获得10
27秒前
马彦娟完成签到 ,获得积分10
30秒前
xuan发布了新的文献求助10
31秒前
白灼虾完成签到 ,获得积分10
31秒前
123完成签到,获得积分10
33秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Römisch-Germanische Forschungen 1000
APA handbook of comparative psychology: Basic concepts, methods, neural substrate, and behavior 1000
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
Electric machines: theory, operating applications, and controls 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7604644
求助须知:如何正确求助?哪些是违规求助? 9180555
关于积分的说明 19661724
捐赠科研通 7179720
什么是DOI,文献DOI怎么找? 3269423
关于科研通互助平台的介绍 2433396
邀请新用户注册赠送积分活动 2263463