计算机科学
杠杆(统计)
行人
稳健性(进化)
人工智能
情报检索
语义学(计算机科学)
几何数据分析
灵活性(工程)
机器学习
语义映射
面子(社会学概念)
人机交互
视频检索
文本检测
图像检索
计算机视觉
自然语言处理
作者
Fanzhi Jiang,Kexin Wang,Hanchi Ren,Y. Li,Liumei Zhang,Yuanjiao Hu,Xianghua Xie,Su Yang
标识
DOI:10.1016/j.engappai.2026.114224
摘要
Person Re-identification (Re-ID) is crucial in computer vision, widely applied in forensic investigation, intelligent surveillance, and video retrieval. Recent text-based Re-ID methods leverage eyewitness descriptions to enhance retrieval flexibility but still face challenges in accurately characterizing individuals under complex conditions. To address issues like low resolution, viewpoint variations, and occlusions, this paper proposes a novel text-based person Re-ID approach that integrates textual descriptions with synthesized Three-dimensional (3D) geometric pedestrian data derived from existing Two-dimensional (2D) images. Specifically, the semantic richness of text compensates for the lack of color and texture details in 3D data, while the robustness of geometric and pose information significantly enhances retrieval performance. Despite current 3D pedestrian data being generated through reconstruction algorithms, this work serves as a pioneering exploration of text-to-3D pedestrian retrieval, offering substantial potential for real-world applications in multimodal biometrics, forensic investigations, and privacy protection. Experiments on three public datasets demonstrate that our method achieves competitive performance, confirming its practical applicability and significance. • First text-based 3D person Re-ID using text and 3D geometry. • Novel framework uses text semantics and 3D geometry to handle occlusions. • Experiments on three datasets show competitive real-world performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI