点云
融合
计算机科学
云计算
人工智能
计算机视觉
操作系统
语言学
哲学
作者
Yaoming Zhuang,Zhenjie Duan,Li Li,Chengdong Wu,Zhanlin Liu
标识
DOI:10.1109/tiv.2024.3513401
摘要
High-definition (HD) maps are crucial for autonomous driving, supporting decision-making, control, and localization. However, current map generation methods often rely on single-modality images, which lack depth information and environmental context. To overcome these limitations, we propose PC-FusionMap, a novel approach that generates point cloud modality data and fuses multi-modal and temporal features for accurate and efficient online HD map construction. Our method addresses the shortcomings of existing approaches by leveraging an improved depth estimation module and a supervised labeling strategy to generate point cloud data. We also introduce a multi-modal feature fusion architecture (CFMB) and a temporal fusion network (TFC) to effectively integrate multi-modal and temporal information. The CFMB architecture uses a query mechanism and cross-attention to enhance the complementary performance between modalities, simplifying the fusion process and improving accuracy. The TFC Network models dynamic changes in time series data, further enhancing the accuracy and robustness of online HD map construction. Our approach achieves state-of-the-art results on the nuScenes and the Argoverse2 datasets, surpassing baseline models in both accuracy and stability. Additionally, incorporating our method into existing HD map generation models can lead to substantial performance gains.
科研通智能强力驱动
Strongly Powered by AbleSci AI