模态(人机交互)
计算机科学
人工智能
特征(语言学)
模式
突出
水准点(测量)
模式识别(心理学)
对象(语法)
目标检测
集合(抽象数据类型)
特征提取
传感器融合
判别式
一般化
融合
计算机视觉
人工神经网络
钥匙(锁)
维数(图论)
匹配(统计)
机器学习
异常
循环神经网络
任务(项目管理)
相似性(几何)
作者
Yang Yang,Nianchang Huang,Qiang Zhang,Jungong Han,J. Huang
标识
DOI:10.1109/tmm.2026.3651119
摘要
This paper delves into the task of arbitrary modality salient object detection (AM SOD), aiming to detect salient objects from the images with arbitrary modality types or arbitrary modality numbers by using a single model trained once. Specifically, we develop a novel model, termed modality adaptive network (MAN), for AM SOD, which addresses two fundamental challenges in AM SOD: the diverse modality discrepancies arising from varying modality types and the dynamic fusion dilemma resulting from an unfixed number of modalities in the input data. Technically, MAN first introduces a novel Modality-Adaptive Feature Extractor (MAFE) to adaptively extract features from different input modalities based on their characteristics by utilizing a set of learnable modality prompts. Concurrently, a new modality translation contractive (MTC) loss is devised to facilitate the training of MAFE as well as modality prompts, thereby effectively addressing the inherent modality discrepancies and extracting more discriminative features from each modality image. Subsequently, MAN presents a hybrid dynamic fusion (HDF) strategy to effectively resolve the challenge of dynamic inputs in multi-modal feature fusion as well as enhance the exploitation of complementary information across different modalities. This is specially achieved by a Channel- wise Dynamic Fusion Module (CDFM) and a Spatial- wise Dynamic Fusion Module (SDFM). Experimental results show that by virtue of MAFE, MTC loss and HDF strategy, our proposed method achieves significant increasements over existing models on benchmark datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI