BEV-CMHF: A Cross-Modality Hybrid Fusion Framework for BEV 3D Object Detection With Feature Interaction and Temporal Fusion

计算机科学 水准点(测量) 人工智能 特征(语言学) 目标检测 推论 保险丝(电气) 计算机视觉 激光雷达 传感器融合 对象(语法) 特征提取 融合 利用 融合机制 高级驾驶员辅助系统 混合动力系统 可视化 领域(数学) 模式识别(心理学) 特征向量 特征检测(计算机视觉)
作者
Jiafeng Li,Jiayi Xu,Mengxun Zhi,Jing Zhang,Li Zhuo
出处
期刊:IEEE Transactions on Intelligent Transportation Systems [Institute of Electrical and Electronics Engineers]
卷期号:27 (5): 5629-5641
标识
DOI:10.1109/tits.2026.3651793
摘要

Autonomous driving technology has garnered significant attention for its potential to reduce driver burden and enhance road safety. Modern autonomous driving systems rely on a variety of sensors to perceive complex driving environments. Many existing methods map heterogeneous data into the bird’s eye view (BEV) space for feature fusion. However, they often fail to fully exploit the cross-modal interactions between cameras and LiDAR, or incorporate temporal information, resulting in suboptimal performance. Furthermore, commonly used fusion strategies are often overly simplistic. This study proposes BEV-CMHF, a cross-modality hybrid fusion framework for BEV 3D object detection with feature interaction and temporal fusion. By introducing an interactive cross-attention module and a long-short-term temporal module, the proposed framework enhances the representational power of fused BEV features. Specifically, a feature-interaction attention module that facilitates effective interaction between the camera and LiDAR BEV features using deformable attention is designed, providing guidance and supervision for the camera BEV features. Subsequently, a historical feature temporal fusion module that integrates the long-short-term temporal module is introduced to incorporate additional critical temporal information into the BEV features. Moreover, a dynamic hybrid feature-fusion module is designed to fuse the BEV features of the camera and LiDAR effectively through a hybrid attention mechanism that combines coarse and fine attention. Extensive experiments conducted on the nuScenes benchmark validate the effectiveness of the proposed method, achieving 70.87% mAP and 74.00% NDS on the test set. Using a single NVIDIA GeForce RTX 4090, the method attained an inference speed of 5.79 images per second (5.79 img/s), corresponding to an inference time of 172.64 ms on the nuScenes dataset. The source code will be released at https://github.com/BJUTsipl/BEV-CMHF
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
1秒前
qpisuo发布了新的文献求助10
1秒前
www发布了新的文献求助10
2秒前
2秒前
兮沐发布了新的文献求助10
2秒前
2秒前
彭于晏应助踏实的十八采纳,获得10
2秒前
12345完成签到,获得积分10
3秒前
gushaohua007关注了科研通微信公众号
3秒前
chen发布了新的文献求助30
4秒前
sunfenghong发布了新的文献求助10
4秒前
5秒前
5秒前
6秒前
大模型应助坚定的贞采纳,获得10
7秒前
7秒前
8秒前
wc发布了新的文献求助10
8秒前
自由飞翔发布了新的文献求助10
8秒前
木易发布了新的文献求助10
9秒前
铁光完成签到,获得积分10
9秒前
WJY完成签到,获得积分10
10秒前
styrene发布了新的文献求助100
10秒前
kimyb发布了新的文献求助10
11秒前
11秒前
好运大王发布了新的文献求助10
12秒前
上官若男应助博士后采纳,获得10
12秒前
科研通AI6.3应助maomao采纳,获得10
12秒前
12秒前
在水一方应助www采纳,获得10
12秒前
细腻荔枝完成签到 ,获得积分10
13秒前
CodeCraft应助杨榆藤采纳,获得10
14秒前
汉堡包应助机灵小蕊采纳,获得10
14秒前
15秒前
16秒前
16秒前
CCL完成签到,获得积分10
17秒前
雪霁凝泫发布了新的文献求助10
18秒前
好运大王完成签到,获得积分10
18秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Resistance Spot Welding Dataset for Automobile Body-in-White Quality Analysis 748
日本現代怪異事典 副読本 700
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 650
Machine Learning for Asset Management and Pricing 600
Numerical analysis of the coupled atmosphere-ocean models (CAO II). II 600
Models for the coupled atmosphere and ocean 600
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7389235
求助须知:如何正确求助?哪些是违规求助? 8995655
关于积分的说明 19143608
捐赠科研通 7026152
什么是DOI,文献DOI怎么找? 3228619
关于科研通互助平台的介绍 2390917
邀请新用户注册赠送积分活动 2209922