计算机科学
人工智能
心理学
反射(计算机编程)
认知心理学
人机交互
可视化
多媒体
自然语言处理
鉴定(生物学)
作者
Sichen Tao,Shuang Li,Jun Ye,Neng Dong,Fan Li,Huafeng Li
标识
DOI:10.1109/tcsvt.2026.3670874
摘要
Video-based Visible-Infrared Person Re-Identification (VVI-ReID) aims to learn consistent person feature representations across video sequences in different modalities. Existing methods that use an intermediate modality to bridge the gap between visible (RGB) and infrared (IR) sequences tend to be limited by high construction costs, loss of high-frequency details, and lack of temporal cues. Moreover, they typically focus on refining global representation using high-level features, neglecting the enhancement of local details through low-level features. To address these challenges, we propose the novel Spatial-Temporal High-Frequency Learning (STHF) framework, which constructs an appropriate intermediate modality for the VVI-ReID task and alleviates the modality gap via hierarchical feature enhancement. Specifically, we introduce the Spatial-Temporal High-Pass Filter (ST-HPF), which filters out spatial-temporal Low-Frequency Components (LFC), preserving high-frequency details to construct an intermediate modality at the sequence level. We then enhance the local details with low-level features through the Shallow Detail Compensation (SDC) module, which reduces local noise interference. Finally, the Deep Semantic Refinement (DSR) module refines the global representation by modeling spatial-temporal high-frequency semantic associations using high-level features. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches on the publicly available HITSZ-VCM and BUPTCampus datasets. The code is available at https://github.com/TSC95720/STHF.
科研通智能强力驱动
Strongly Powered by AbleSci AI