适配器(计算)
计算机科学
情态动词
人工智能
蒸馏
视频跟踪
计算机视觉
跟踪(教育)
眼动
对象(语法)
计算机硬件
化学
色谱法
心理学
教育学
高分子化学
作者
Xiaojun Hou,Jiazheng Xing,Yijie Qian,Yaowei Guo,Shuo Xin,Junhao Chen,Kai Tang,Rui Wang,Zhengkai Jiang,Liang Liu,Yong Liu
出处
期刊:
日期:2024-06-16
卷期号:: 26541-26551
被引量:74
标识
DOI:10.1109/cvpr52733.2024.02507
摘要
Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers, which was inefficient and lacked generalized representation due to the scarcity of multimodal data. Therefore, recent studies have utilized prompt tuning to transfer pre-trained RGB-based trackers to multimodal data. However, the modality gap limits pre-trained knowledge recall, and the dominance of the RGB modality persists, preventing the full utilization of information from other modalities. To address these issues, we propose a novel symmetric multimodal tracking framework called SDSTrack. We introduce lightweight adaptation for efficient fine-tuning, which directly transfers the feature extraction ability from RGB to other domains with a small number of trainable parameters and integrates multimodal features in a balanced, symmetric manner. Furthermore, we design a complementary masked patch distillation strategy to enhance the robustness of trackers in complex environments, such as extreme weather, poor imaging, and sensor failure. Extensive experiments demonstrate that SDSTrack outperforms state-of-the-art methods in various multimodal tracking scenarios, including RGB+Depth, RGB+Thermal, and RGB+Event tracking, and exhibits impressive results in extreme conditions. Our source code is available at: https://github.com/hoqolo/SDSTrack.
科研通智能强力驱动
Strongly Powered by AbleSci AI