计算机科学
水准点(测量)
人工智能
特征(语言学)
目标检测
推论
保险丝(电气)
计算机视觉
激光雷达
传感器融合
对象(语法)
特征提取
融合
利用
融合机制
高级驾驶员辅助系统
混合动力系统
可视化
领域(数学)
模式识别(心理学)
特征向量
特征检测(计算机视觉)
作者
Jiafeng Li,Jiayi Xu,Mengxun Zhi,Jing Zhang,Li Zhuo
标识
DOI:10.1109/tits.2026.3651793
摘要
Autonomous driving technology has garnered significant attention for its potential to reduce driver burden and enhance road safety. Modern autonomous driving systems rely on a variety of sensors to perceive complex driving environments. Many existing methods map heterogeneous data into the bird’s eye view (BEV) space for feature fusion. However, they often fail to fully exploit the cross-modal interactions between cameras and LiDAR, or incorporate temporal information, resulting in suboptimal performance. Furthermore, commonly used fusion strategies are often overly simplistic. This study proposes BEV-CMHF, a cross-modality hybrid fusion framework for BEV 3D object detection with feature interaction and temporal fusion. By introducing an interactive cross-attention module and a long-short-term temporal module, the proposed framework enhances the representational power of fused BEV features. Specifically, a feature-interaction attention module that facilitates effective interaction between the camera and LiDAR BEV features using deformable attention is designed, providing guidance and supervision for the camera BEV features. Subsequently, a historical feature temporal fusion module that integrates the long-short-term temporal module is introduced to incorporate additional critical temporal information into the BEV features. Moreover, a dynamic hybrid feature-fusion module is designed to fuse the BEV features of the camera and LiDAR effectively through a hybrid attention mechanism that combines coarse and fine attention. Extensive experiments conducted on the nuScenes benchmark validate the effectiveness of the proposed method, achieving 70.87% mAP and 74.00% NDS on the test set. Using a single NVIDIA GeForce RTX 4090, the method attained an inference speed of 5.79 images per second (5.79 img/s), corresponding to an inference time of 172.64 ms on the nuScenes dataset. The source code will be released at https://github.com/BJUTsipl/BEV-CMHF
科研通智能强力驱动
Strongly Powered by AbleSci AI