Image-text multimodal classification via cross-attention contextual transformer with modality-collaborative learning

模式 模态(人机交互) 计算机科学 多模式学习 人工智能 机器学习 传感器融合 社会科学 社会学
作者
Qianyao Shi,Wanru Xu,Zhenjiang Miao
出处
期刊:Journal of Electronic Imaging [SPIE]
卷期号:33 (04)
标识
DOI:10.1117/1.jei.33.4.043042
摘要

Nowadays, we are surrounded by various types of data from different modalities, such as text, images, audio, and video. The existence of this multimodal data provides us with rich information, but it also brings new challenges: how do we effectively utilize this data for accurate classification? This is the main problem faced by multimodal classification tasks. Multimodal classification is an important task that aims to classify data from different modalities. However, due to the different characteristics and structures of data from different modalities, effectively fusing and utilizing them for classification is a challenging problem. To address this issue, we propose a cross-attention contextual transformer with modality-collaborative learning for multimodal classification (CACT-MCL-MMC) to better integrate information from different modalities. On the one hand, existing multimodal fusion methods ignore the intra- and inter-modality relationships, and there is unnoticed information in the modalities, resulting in unsatisfactory classification performance. To address the problem of insufficient interaction of modality information in existing algorithms, we use a cross-attention contextual transformer to capture the contextual relationships within and among modalities to improve the representativeness of the model. On the other hand, due to differences in the quality of information among different modalities, some modalities may have misleading or ambiguous information. Treating each modality equally may result in modality perceptual noise, which reduces the performance of multimodal classification. Therefore, we use modality-collaborative to filter misleading information, alleviate the quality difference of information among modalities, align modality information with high-quality and effective modalities, enhance unimodal information, and obtain more ideal multimodal fusion information to improve the model's discriminative ability. Our comparative experimental results on two benchmark datasets for image-text classification, CrisisMMD and UPMC Food-101, show that our proposed model outperforms other classification methods and even state-of-the-art (SOTA) multimodal classification methods. Meanwhile, the effectiveness of the cross-attention module, multimodal contextual attention network, and modality-collaborative learning was verified through ablation experiments. In addition, conducting hyper-parameter validation experiments showed that different fusion calculation methods resulted in differences in experimental results. The most effective feature tensor calculation method was found. We also conducted qualitative experiments. Compared with the original model, our proposed model can identify the expected results in the vast majority of cases. The codes are available at https://github.com/KobeBryant8-24-MVP/CACT-MCL-MMC. The CrisisMMD is available at https://dataverse.mpisws.org/dataverse/icwsm18, and the UPMC-Food-101 is available at https://visiir.isir.upmc.fr/.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
齐朕完成签到,获得积分10
1秒前
Pendulium发布了新的文献求助10
1秒前
Akim应助自由的聋五采纳,获得10
1秒前
骤世界完成签到 ,获得积分10
1秒前
2秒前
整齐的灵雁完成签到,获得积分10
2秒前
Amber完成签到,获得积分20
2秒前
bkagyin应助Gloria采纳,获得10
2秒前
2秒前
笔笔完成签到,获得积分10
2秒前
一7完成签到,获得积分10
2秒前
Lucas应助追寻紫安采纳,获得10
3秒前
3秒前
六氟合铂酸氙完成签到,获得积分10
3秒前
RoseSpire完成签到,获得积分10
3秒前
HHHHTTTT应助乐之采纳,获得10
3秒前
大模型应助苍禾采纳,获得30
4秒前
科研通AI6.3应助蓝天采纳,获得30
4秒前
无情采文发布了新的文献求助10
4秒前
天天快乐应助复杂采纳,获得10
4秒前
Yonina发布了新的文献求助10
5秒前
yyj发布了新的文献求助10
5秒前
GAC发布了新的文献求助30
6秒前
科研通AI6.4应助peACE采纳,获得10
6秒前
7秒前
7秒前
Yanxb发布了新的文献求助10
7秒前
7秒前
英俊的铭应助Amber采纳,获得10
7秒前
英姑应助qqqqqxing采纳,获得10
8秒前
是小妤呀完成签到 ,获得积分10
8秒前
Hello应助shxygpz采纳,获得10
8秒前
鸡蛋酱完成签到 ,获得积分10
8秒前
8秒前
8秒前
汉堡包应助HCS采纳,获得10
9秒前
NexusExplorer应助路途中追逐采纳,获得10
10秒前
10秒前
10秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Nondestructive Testing Handbook: Vol. 4, Thermal and Infrared Testing (IR), 4th ed 800
作者名:Kristopher P. Plain,悉尼大学的,目前只能查到其四篇论文,想找到其博士论文 590
Évora na Idade Média 555
Soil mites of the family Rhagidiidae (Actinedida: Eupodoidea). Morphology, Systematics, Ecology 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Radical Reactions 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7356758
求助须知:如何正确求助?哪些是违规求助? 8967456
关于积分的说明 19054532
捐赠科研通 7004380
什么是DOI,文献DOI怎么找? 3222316
关于科研通互助平台的介绍 2386476
邀请新用户注册赠送积分活动 2202905