计算机科学
编码器
人工智能
跟踪(教育)
一般化
计算机视觉
RGB颜色模型
视频跟踪
运动(物理)
实时计算
跟踪系统
帧速率
空间智能
深度学习
模式识别(心理学)
传感器融合
匹配移动
机器学习
作者
Hangfei Li,G.X. Liu,Xiao Guo,Yufei Zha,Peng Zhang
标识
DOI:10.1109/tmm.2026.3664973
摘要
This paper presents a novel video-level RGB-T tracking paradigm based on prompt learning, termed PromptTrack, which establishes dense spatial-temporal associations through cross-modal interactions. The method introduces streaming temporal prompts to capture continuous target dynamics (e.g., appearance changes and motion trajectories), while leveraging multimodal spatial prompts to utilize complementary RGB and thermal-infrared (TIR) features dynamically. By propagating temporal prompts through consecutive frames and integrating bidirectional spatial interactions between modalities, PromptTrack achieves superior tracking performance in complex scenarios such as occlusion, low illumination, and distractors. The proposed framework employs a unified multimodal encoder with spatial-temporal modeling via multimodal spatial prompt blocks, enabling efficient fusion of RGB-TIR features without requiring domain-specific structure modifications. Extensive experiments on three RGB-T benchmarks (LasHeR, RGBT210, RGBT234) demonstrate that PromptTrack achieves new state-of-the-art performance, with 76.2% in precision rate on LasHeR and outperforming existing methods by ++1.9% in precision rate and +0.5% in success rate on RGBT234. Notably, its modality-agnostic design facilitates seamless generalization to RGB-D and RGB-E tracking domains achieving new benchmarks on DepthTrack, VOT-RGBD2022, and VisEvent datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI