Cross-Modality Spatial-Temporal Transformer for Video-Based Visible-Infrared Person Re-Identification

计算机科学 人工智能 模式识别(心理学) 计算机视觉 变压器 特征学习 相互信息 空间分析 数学 统计 物理 量子力学 电压
作者
Yujian Feng,Feng Chen,Jian Yu,Yimu Ji,Fei Wu,Tianliang Liu,Shangdong Liu,Xiao‐Yuan Jing,Jiebo Luo
出处
期刊:IEEE Transactions on Multimedia [Institute of Electrical and Electronics Engineers]
卷期号:26: 6582-6594 被引量:4
标识
DOI:10.1109/tmm.2024.3354575
摘要

Video-based visible-infrared person re-identification (VVI-ReID) aims to match the identity of a person captured in video sequences from both visible and infrared cameras. The VVI-ReID task requires considering both the spatial relationship between body parts within each frame and the temporal change of appearance between successive frames. Existing VVI Re-ID methods employ Convolutional Neural Networks to extract local spatial features and Long Short-Term Memory to form temporal associations. However, these methods can not effectively capture the global spatial feature and the long-range temporal dependencies in ultra-long sequences. In this paper, we propose a Cross-modality Spatial-temporal Transformer (CST) including a Cross-frame Tube Transformer Module (CTTM) and a Multi-frame Transformer Fusion Module (MTFM) to address these challenges. Firstly, CTTM tokenizes a video clip into multiple 3D tubes, each encapsulating local spatial-temporal information of pedestrians, and then obtains global spatial-temporal representations by establishing the relationship between tubes. Secondly, we design MTFM to exchange information between multiple frames using message tokens, thus modeling the long-range temporal dependencies of features of pedestrians. In addition, to prevent the potential representation collapse caused by triplet-based loss functions, we propose a diversity-consistency (DC) loss function to preserve the diversity and consistency of cross-modality feature representations by imposing variance, invariance, and covariance constraints in feature representations. Extensive benchmark experiments demonstrate that our approach outperforms the state-of-the-art methods with large margins.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
科研通AI6.4应助buerxiaoshen采纳,获得10
刚刚
Nole应助踏实书白采纳,获得10
刚刚
刚刚
桐桐应助真一松采纳,获得10
刚刚
刚刚
不会游泳的鱼完成签到,获得积分10
1秒前
heher完成签到 ,获得积分0
2秒前
Hantheex发布了新的文献求助10
2秒前
DAYDAY发布了新的文献求助10
3秒前
tsd完成签到,获得积分10
3秒前
3秒前
迷人沛儿发布了新的文献求助10
3秒前
多来A梦发布了新的文献求助10
4秒前
4秒前
筱汐完成签到,获得积分10
4秒前
毕业顺利发布了新的文献求助10
5秒前
6秒前
领导范儿应助大米饭采纳,获得10
6秒前
7秒前
慕青应助yorkson境采纳,获得10
7秒前
情怀应助袁袁袁采纳,获得30
8秒前
8秒前
8秒前
李健应助甜蜜花采纳,获得10
8秒前
8秒前
陈住气发布了新的文献求助10
8秒前
9秒前
木木发布了新的文献求助10
9秒前
9秒前
Hantheex发布了新的文献求助10
11秒前
WJY发布了新的文献求助30
11秒前
小鱼儿发布了新的文献求助10
11秒前
Lucas应助披萨好吃酱采纳,获得10
12秒前
英俊的铭应助Xxxx采纳,获得10
13秒前
蓝天发布了新的文献求助10
13秒前
bkagyin应助易酰水烊酸采纳,获得10
13秒前
13秒前
HY完成签到 ,获得积分10
14秒前
kleinxiao发布了新的文献求助10
14秒前
baobao发布了新的文献求助10
14秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Atlas of Aligner Treatment and Planning A Case-Based Approach 1000
Geist der Kunst und Kultur 1000
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 700
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
丝光沸石活性位点定向调控及其二甲醚羰基化性能研究 500
A Concise History of the World, 2nd Edition 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7420709
求助须知:如何正确求助?哪些是违规求助? 9024200
关于积分的说明 19224077
捐赠科研通 7051017
什么是DOI,文献DOI怎么找? 3235024
关于科研通互助平台的介绍 2397978
邀请新用户注册赠送积分活动 2217228