A survey of multimodal hybrid deep learning for computer vision: Architectures, applications, trends, and challenges

深度学习 人工智能 计算机科学 机器学习 领域(数学) 模式 多模式学习 社会科学 数学 社会学 纯数学
作者
Khaled Bayoudh
出处
期刊:Information Fusion [Elsevier BV]
卷期号:105: 102217-102217 被引量:85
标识
DOI:10.1016/j.inffus.2023.102217
摘要

In recent years, deep learning algorithms have rapidly revolutionized artificial intelligence, particularly machine learning, enabling researchers and practitioners to extend previously hand-crafted feature extraction procedures. In particular, deep learning uses adaptive learning processes to learn more complex and informative patterns from datasets of varying sizes. With the increasing availability of multimodal data streams and recent advances in deep learning algorithms, multimodal deep learning is on the rise. This requires the development of complex models that can process and analyze multimodal information in a consistent manner. However, unstructured data can come in many different forms (also known as modalities). Extracting relevant features from this data remains an ambitious goal for deep learning researchers. According to the literature, most deep learning systems consist of a single architecture (i.e., standalone deep learning). When two or more deep learning architectures are combined over multiple sensory modalities, the result is called a multimodal hybrid deep learning model. Since this research direction has received much attention in the field of deep learning, the purpose of this survey is to provide a broader overview of the topic. In this paper, we provide a comprehensive review of recent advances in multimodal hybrid deep learning, including a thorough analysis of the most commonly developed hybrid architectures. In particular, one of the main challenges in multimodal hybrid analysis is the ability of these architectures to systematically integrate cross-modal features in hybrid designs. Therefore, we propose a generic framework for multimodal hybrid learning that focuses mainly on fusion methods. We also identify trends and challenges in multimodal hybrid learning and provide insights and directions for future research. Our findings show that multimodal hybrid learning can perform well in a variety of challenging computer vision applications and tasks.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
科研通AI6.4应助危机的囧采纳,获得30
刚刚
陆文灏发布了新的文献求助10
刚刚
lee完成签到 ,获得积分10
刚刚
1秒前
鲍里斯瓦格完成签到,获得积分10
2秒前
一一发布了新的文献求助10
2秒前
锌银12306发布了新的文献求助10
3秒前
陆文灏发布了新的文献求助10
3秒前
wanci应助美满的无极采纳,获得10
3秒前
陆文灏发布了新的文献求助10
3秒前
陆文灏发布了新的文献求助10
3秒前
4秒前
扬大小汤发布了新的文献求助10
4秒前
4秒前
5秒前
5秒前
5秒前
5秒前
6秒前
陆文灏发布了新的文献求助10
6秒前
科研通AI6.4应助向北游采纳,获得10
8秒前
Why完成签到,获得积分10
8秒前
消逝发布了新的文献求助10
9秒前
qumingzihaonan完成签到,获得积分10
9秒前
陆文灏发布了新的文献求助10
9秒前
陆文灏发布了新的文献求助10
10秒前
陆文灏发布了新的文献求助10
10秒前
10秒前
aaa发布了新的文献求助10
11秒前
美好的冰蓝完成签到 ,获得积分10
11秒前
12秒前
锌银12306完成签到,获得积分10
12秒前
12秒前
可爱的函函应助扬大小汤采纳,获得10
12秒前
多情无敌完成签到 ,获得积分10
13秒前
我是老大应助Neptune采纳,获得10
13秒前
CodeCraft应助着急的听枫采纳,获得10
14秒前
xudonghui发布了新的文献求助10
16秒前
18秒前
向北游发布了新的文献求助10
18秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
2026年中国辛酸癸酸聚乙二醇甘油酯行业市场现状调查及投资机会研判报告 1000
模型平均及其应用 900
Nondestructive Testing Handbook: Vol. 4, Thermal and Infrared Testing (IR), 4th ed 800
Évora na Idade Média 555
作者名:Kristopher P. Plain,悉尼大学的,目前只能查到其四篇论文,想找到其博士论文 550
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7345593
求助须知:如何正确求助?哪些是违规求助? 8957862
关于积分的说明 19022018
捐赠科研通 6996873
什么是DOI,文献DOI怎么找? 3220000
关于科研通互助平台的介绍 2384901
邀请新用户注册赠送积分活动 2200252