mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

计算机科学 杠杆(统计) 情态动词 人工智能 模式 对话 自然语言处理 语言学 社会科学 哲学 社会学 化学 高分子化学
作者
Qinghao Ye,Haiyang Xu,Guohai Xu,Jiabo Ye,Ming Yan,Yiyang Zhou,Junyang Wang,Anwen Hu,Pengcheng Shi,Yaya Shi,Chenliang Li,Yuanhong Xu,Hehong Chen,Junfeng Tian,Qian Qi,Ji Zhang,Fei Huang
出处
期刊:Cornell University - arXiv [Cornell University]
被引量:133
标识
DOI:10.48550/arxiv.2304.14178
摘要

Large language models (LLMs) have demonstrated impressive zero-shot abilities on a variety of open-ended tasks, while recent research has also explored the use of LLMs for multi-modal generation. In this study, we introduce mPLUG-Owl, a novel training paradigm that equips LLMs with multi-modal abilities through modularized learning of foundation LLM, a visual knowledge module, and a visual abstractor module. This approach can support multiple modalities and facilitate diverse unimodal and multimodal abilities through modality collaboration. The training paradigm of mPLUG-Owl involves a two-stage method for aligning image and text, which learns visual knowledge with the assistance of LLM while maintaining and even improving the generation abilities of LLM. In the first stage, the visual knowledge module and abstractor module are trained with a frozen LLM module to align the image and text. In the second stage, language-only and multi-modal supervised datasets are used to jointly fine-tune a low-rank adaption (LoRA) module on LLM and the abstractor module by freezing the visual knowledge module. We carefully build a visually-related instruction evaluation set OwlEval. Experimental results show that our model outperforms existing multi-modal models, demonstrating mPLUG-Owl's impressive instruction and visual understanding ability, multi-turn conversation ability, and knowledge reasoning ability. Besides, we observe some unexpected and exciting abilities such as multi-image correlation and scene text understanding, which makes it possible to leverage it for harder real scenarios, such as vision-only document comprehension. Our code, pre-trained model, instruction-tuned models, and evaluation set are available at https://github.com/X-PLUG/mPLUG-Owl. The online demo is available at https://www.modelscope.cn/studios/damo/mPLUG-Owl.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
搬砖吗喽完成签到,获得积分10
4秒前
nanfeng完成签到 ,获得积分10
4秒前
LLin完成签到,获得积分10
5秒前
思念变成王年年完成签到,获得积分10
8秒前
自由念露完成签到 ,获得积分10
9秒前
ly完成签到,获得积分10
11秒前
嘻嘻我完成签到,获得积分10
12秒前
CWC完成签到,获得积分10
12秒前
wzhang完成签到,获得积分10
13秒前
孝顺的天思完成签到 ,获得积分10
14秒前
wing完成签到 ,获得积分10
17秒前
Faye完成签到 ,获得积分10
29秒前
JUAN完成签到,获得积分10
36秒前
任志政完成签到 ,获得积分10
39秒前
飞快的蛋完成签到,获得积分0
40秒前
嘟嘟完成签到 ,获得积分10
41秒前
整齐听南完成签到 ,获得积分10
42秒前
run完成签到 ,获得积分10
42秒前
李爱国应助科研通管家采纳,获得10
45秒前
Akim应助科研通管家采纳,获得10
45秒前
Lucas应助科研通管家采纳,获得20
45秒前
lzj完成签到,获得积分10
45秒前
cdercder应助科研通管家采纳,获得10
45秒前
cdercder应助科研通管家采纳,获得10
45秒前
45秒前
cdercder应助科研通管家采纳,获得10
46秒前
cdercder应助科研通管家采纳,获得20
46秒前
张匀继完成签到 ,获得积分10
46秒前
田様应助李寳采纳,获得10
46秒前
TT完成签到 ,获得积分20
48秒前
微卫星不稳定完成签到 ,获得积分0
49秒前
义气的惜霜完成签到 ,获得积分10
50秒前
Joy完成签到,获得积分10
52秒前
czxcpu完成签到 ,获得积分10
54秒前
隐形的小鸽子完成签到 ,获得积分10
55秒前
粗心的逍遥完成签到 ,获得积分10
55秒前
56秒前
weng完成签到,获得积分10
56秒前
xingyuwuhen007完成签到,获得积分10
59秒前
Jackson发布了新的文献求助30
1分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
化工安全与环保 1000
Autoparametric Resonance in Mechanical Systems 1000
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 800
Effects of Two Weeks of Red Light Therapy on Choroidal Thickness and Axial Length in Young Adults 700
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7656682
求助须知:如何正确求助?哪些是违规求助? 9227336
关于积分的说明 19829001
捐赠科研通 7223111
什么是DOI,文献DOI怎么找? 3280333
关于科研通互助平台的介绍 2440621
邀请新用户注册赠送积分活动 2280175