判别式
计算机科学
互惠的
特征学习
人工智能
代表(政治)
特征(语言学)
模式识别(心理学)
图像检索
方案(数学)
粒度
机器学习
编码
任务(项目管理)
特征提取
深度学习
人工神经网络
特征向量
成对比较
任务分析
鉴定(生物学)
答疑
模态(人机交互)
图像(数学)
视觉文字
可扩展性
语义学(计算机科学)
作者
Anh D. Nguyen,Hoa N. Nguyen
标识
DOI:10.1109/tip.2025.3594880
摘要
Text-based person retrieval is defined as the challenging task of searching for people's images based on given textual queries in natural language. Conventional methods primarily use deep neural networks to understand the relationship between visual and textual data, creating a shared feature space for cross-modal matching. The absence of awareness regarding variations in feature granularity between the two modalities, coupled with the diverse poses and viewing angles of images corresponding to the same individual, may lead to overlooking significant differences within each modality and across modalities, despite notable enhancements. Furthermore, the inconsistency in caption queries in large public datasets presents an additional obstacle to cross-modality mapping learning. Therefore, we introduce 3RTPR, a novel text-based person retrieval method that integrates a representation fusing mechanism and an adaptive loss refinement algorithm into a dual-encoder branch architecture. Moreover, we propose training two independent models simultaneously, which reciprocally support each other to enhance learning effectiveness. Consequently, our approach encompasses three significant contributions: (i) proposing a fused representation method to generate more discriminative representations for images and captions; (ii) introducing a novel algorithm to adjust loss and prioritize samples that contain valuable information; and (iii) proposing reciprocal learning involving a pair of independent models, which allows us to enhance general retrieval performance. In order to validate our method's effectiveness, we also demonstrate superior performance over state-of-the-art methods by performing rigorous experiments on three well-known benchmarks: CUHK-PEDES, ICFG-PEDES, and RSTPReid.
科研通智能强力驱动
Strongly Powered by AbleSci AI