计算机科学
编码器
变压器
模式
差异(会计)
粒度
人工智能
模态(人机交互)
匹配(统计)
编码(集合论)
图像(数学)
班级(哲学)
模式识别(心理学)
情报检索
机器学习
数据挖掘
集合(抽象数据类型)
程序设计语言
社会学
电压
业务
会计
物理
操作系统
统计
量子力学
社会科学
数学
作者
Liping Bao,Longhui Wei,Wengang Zhou,Lin Liu,Lingxi Xie,Houqiang Li,Qi Tian
标识
DOI:10.1109/tmm.2023.3321504
摘要
Text-based person search aims to retrieve the most relevant pedestrian images from an image gallery based on textual descriptions. Most existing methods rely on two separate encoders to extract the image and text features, and then elaborately design various schemes to bridge the gap between image and text modalities. However, the shallow interaction between both modalities in these methods is still insufficient to eliminate the modality gap. To address the above problem, we propose TransTPS, a transformer-based framework that enables deeper interaction between both modalities through the self-attention mechanism in transformer, effectively alleviating the modality gap. In addition, due to the small inter-class variance and large intra-class variance in image modality, we further develop two techniques to overcome these limitations. Specifically, Cross-modal Multi-Granularity Matching (CMGM) is proposed to address the problem caused by small inter-class variance and facilitate distinguishing pedestrians with similar appearance. Besides, Contrastive Loss with Weakly Positive pairs (CLWP) is introduced to mitigate the impact of large intra-class variance and contribute to the retrieval of more target images. Experiments on CUHK-PEDES and RSTPReID datasets demonstrate that our proposed framework achieves state-of-the-art performance compared to previous methods. Our code will be available at https://github.com/baolp/TransTPS .
科研通智能强力驱动
Strongly Powered by AbleSci AI