计算机科学
人工智能
特征学习
语义计算
自然语言处理
卷积神经网络
联想(心理学)
图形
语义学(计算机科学)
语义特征
特征(语言学)
一般化
语义记忆
模态(人机交互)
钥匙(锁)
语义相似性
光学(聚焦)
深度学习
代表(政治)
语义压缩
特征选择
特征向量
模式
语义整合
显式语义分析
语义网络
情报检索
语义映射
监督学习
无监督学习
机器学习
语义空间
语义网格
空格(标点符号)
语义数据模型
特征提取
作者
Bin Yang,Lekai Liu,Wenke Huang,Xiao Wang,Bo Du,Mang Ye
标识
DOI:10.1109/tifs.2025.3645635
摘要
Unsupervised visible-infrared person reidentification (US-VI-ReID) seeks to learn a cross-modality retrieval model without relying on manual annotations, thereby reducing the high cost associated with labeling. Recent large-scale vision-language pre-training models, such as CLIP, have shown significant potential in enhancing pure-vision-based person re-identification. However, existing CLIP-based US-VI-ReID methods focus on independently learning semantic information within the visible and infrared modalities. These methods overlook the mismatch between the pre-training data of CLIP and the downstream cross-modality data, resulting in substantial cross-modal semantic differences. Such inconsistent semantic information, which exhibits modality discrepancies, cannot ensure the accuracy of cross-modality associations and thus hampers the performance of cross-modality learning. To address these challenges and further explore the generalizable semantic representation across modalities in CLIP, we propose a novel framework named Mining Cross-Modality Implicit Semantic Association (MCSA), which focuses on learning a modality-invariant implicit semantic space to enhance cross-modality associations and feature learning. The proposed method comprises two key modules: Modality-invariant Prompt Learning and GCNs-Driven Collaboration Alignment. Specifically, to enable CLIP to learn modality-invariant semantics, we integrate a random color augmentation branch into the visible stream for joint contrastive learning for mining generalizable semantic representations. This ensures the color generalization of the constructed implicit semantic prompts. Moreover, within the cross-modal invariant implicit semantic space, we utilize Graph Convolutional Networks (GCNs) to uncover more reliable cross-modal associations. By integrating information from images and semantic graphs, we jointly refine cross-modal correspondences, enabling the model to perform precise cross-modal feature learning. Extensive experiments conducted on the SYSU-MM01 and RegDB datasets demonstrate the effectiveness of the proposed MCSA. The source code will be released.
科研通智能强力驱动
Strongly Powered by AbleSci AI