计算机科学
适配器(计算)
情态动词
人工智能
比例(比率)
融合
传感器融合
机器学习
操作系统
语言学
化学
物理
哲学
量子力学
高分子化学
作者
Xiang Liu,Haiyan Li,Victor S. Sheng,Yujun Ma,Xiaoguo Liang,Guanbo Wang
标识
DOI:10.1109/tmm.2025.3623526
摘要
Fusing visible (RGB) and thermal (T) images for RGBT tracking has received growing interest in the field of computer vision. However, how to improve the robustness of the tracker to target scale variety, effectively apply visual prompts to multimodal tracking tasks, and enhance the multimodal fusion effectiveness are still urgent challenges in the field of RGBT tracking. To this purpose, this work proposes an RGBT tracking framework integrating scale-aware dilation attention, multimodal prompt interaction learning, and cross- fusion adapter, named MPANet. Firstly, a scale-aware dilation attention (SADA) module is put forward to enhance the flexibility of the tracker in the presence of target scale variations by embedding convolutions with different dilation rates into the self-attention. Subsequently, a multimodal prompt interaction learning (MPIL) module is constructed, which combines global token adaptive attention and spatial attention to efficiently learn visual prompts from different modalities and achieve intermodal prompt interactions. Finally, a cross-fusion adapter (CFA) is developed to facilitate the adaptability of the network to different modalities in the process of multimodal information fusion through the adapter mechanism. Extensive experiments on public RGBT benchmark tracking datasets such as GTOT, RGBT234, LasHeR and VTUAV demonstrate that the proposed method outperforms existing advanced trackers and achieves state-of-the-art performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI