代表(政治)
情态动词
计算机科学
体素
对象(语法)
计算机视觉
目标检测
人工智能
模式识别(心理学)
材料科学
高分子化学
政治学
政治
法学
作者
Kaiqi Liu,Yuanyuan Deng,Jiaxun Tong,Wei Li
标识
DOI:10.1109/jsen.2025.3589494
摘要
Fusing camera and LiDAR information is one of the effective means for achieving robust 3D object detection. However, current 3D multi-modal methods typically rely on independent branches to extract features from different sensors separately, leading to underutilization of complementary information. In this paper, a multi-modal detector named UniVoxel is proposed, which is built on a query-based detection paradigm. The UniVoxel integrates inputs from various modalities into the voxel representation for fusion. Specifically, a Semantic-guided Query Generator (SQG) is proposed, in which the low-level voxel features are utilized to adaptively sample multi-scale image features, producing unified multi-modal voxel features. The multi-modal voxel features contain both the geometric and semantic information of the voxels and can ensure that the model focuses on the Regions of Interest (RoI). Meanwhile, for maximizing the utilization of complementary information, a Fusion Voxel Encoder (FVE) is introduced to update the multi-modal voxels through interacting with the multi-scale semantic information of different cameras. Extensive experiments are conducted on the nuScenes dataset. With the help of the proposed framework, the precision of the object detection has been improved both on the validation set and the test set.
科研通智能强力驱动
Strongly Powered by AbleSci AI