计算机科学
情态动词
比例(比率)
传感器融合
人工智能
融合
情报检索
计算机视觉
地理
材料科学
语言学
地图学
哲学
高分子化学
作者
Zhou Wu,Yu Wang,Hongtuo Qi,Liang Feng,Jiepeng Liu
标识
DOI:10.1109/tii.2025.3584548
摘要
2D-3D cross modal retrieval (CMR) aims to retrieve query image matching points from a 3D reference map. Existing classical CMR datasets and methods commonly support database-based retrieval only, i.e., the point cloud retrieval results are fixed-scale geometric surfaces. The failure to consider geometric regions and information scales fundamentally limits the practical deployment of CMR in engineering systems that require dynamic spatial reasoning, such as autonomous navigation or three-dimensional industrial measurement. In this article, we introduce a new benchmark called cross modal regional retrieval, which extends the classic CMR to allow the free retrieval of associated regions within the point cloud from images. Toward this, a multiview training paradigm is proposed in the training phase, which enables the model to identify occluded points in the region based on a single view. Autoencoders are utilized to learn the mapping of fusion features from a single view to multiple views. We also convert the image retrieval task within the scene cloud into a point classification task in the image to implement global free retrieval. The information fusion and guidance provided by the global point cloud enhances the capability of image cross-modal retrieval. To match the input patterns of the model, we propose a method for constructing datasets from three benchmark sources. Extensive experiments demonstrate that our method achieves state-of-the-art performance compared to existing methods for 2D-3D cross modal regional retrieval.
科研通智能强力驱动
Strongly Powered by AbleSci AI