计算机科学
人工智能
计算机视觉
卷积神经网络
传感器融合
情态动词
模态(人机交互)
特征(语言学)
信息融合
人工神经网络
跟踪(教育)
RGB颜色模型
噪音(视频)
视频跟踪
目标检测
融合
特征提取
对象(语法)
变压器
眼动
可视化
模式识别(心理学)
模式
跟踪系统
语义学(计算机科学)
作者
Jun Liu,Wei Ke,Shuai Wang,Da Yang,Hao Sheng
标识
DOI:10.1109/lsp.2025.3638688
摘要
Visual tracking that combines RGB and thermal infrared modalities (RGB-T) aims to utilize the useful information of each modality to achieve more robust object localization. Most existing tracking methods based on convolutional neural networks (CNNs) and Transformers emphasize integrating multi-modal features through cross-modal attention, but ignore the potential exploitability of complementary information learned by cross-modal attention for enhancing modal features. In this paper, we propose a novel hierarchical progressive fusion network based on cross-modal attention guided enhancement for RGB-T tracking. Specifically, the complementary information generated by cross-modal attention implicitly reflects the consistent regions of interest of important information between different modalities, which is used to enhance modal features in a targeted manner. In addition, a modal feature refinement module and a fusion module are designed based on dynamic routing to perform noise suppression and adaptive integration on the enhanced multi-modal features. Extensive experiments on GTOT, RGBT234, LasHeR and VTUAV show that our method has competitive performance compared with recent state-of-the-art methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI