概率逻辑
计算机科学
人工智能
稳健性(进化)
嵌入
模棱两可
模式识别(心理学)
统计模型
概率分布
偏离随机性模型
可视化
代表(政治)
特征(语言学)
计算机视觉
特征提取
高斯分布
图像(数学)
图像检索
任务(项目管理)
钥匙(锁)
行人检测
任务分析
混合模型
自然语言处理
机器学习
行人
概率方法
上下文图像分类
数据挖掘
视觉文字
计算机视觉中的词袋模型
特征向量
作者
Xi Yang,Kun Chen,Chenghuan Qi,Nannan Wang
标识
DOI:10.1109/tcsvt.2026.3662704
摘要
Text-based person retrieval is a cross-modal task that seeks to match pedestrian images with their corresponding textual descriptions. A key challenge in this task arises from the inherent one-to-many relationships: a single image can correspond to multiple descriptions, and a single description may relate to several images. Conventional deterministic embedding methods, which map images and texts to fixed feature vectors, struggle to capture such complex relationships effectively. To overcome this limitation, we introduce Probabilistic Distribution Alignment (PDA), a framework that represents both pedestrian images and text as probabilistic distributions and models the interactions between visual and linguistic modalities. PDA comprises three main components. First, Distributional Representation Modeling (DRM) encodes images and text into Gaussian distributions using a specially designed distance metric, allowing the model to capture uncertainty in the representations. Second, Cross-Modal Containment (CMC) aligns the distributions of text and masked text with their associated image distributions to strengthen semantic correspondence. Third, Intra-Modal Containment (IMC) enforces structured learning within each modality by embedding distributions alongside their masked variants, improving robustness to incomplete observations. Experiments on standard benchmarks demonstrate that PDA achieves superior performance compared with state-of-the-art methods, effectively handling ambiguity and cross-modal variability. These results highlight probabilistic distribution modeling as a powerful paradigm for vision-language alignment in pedestrian retrieval.
科研通智能强力驱动
Strongly Powered by AbleSci AI