SSLFusion: Scale and Space Aligned Latent Fusion Model for Multimodal 3D Object Detection

比例(比率) 空格(标点符号) 计算机科学 对象(语法) 人工智能 缩放空间 计算机视觉 融合 模式识别(心理学) 地理 地图学 图像处理 图像(数学) 语言学 操作系统 哲学
作者
Bonan Ding,Jin Xie,Jing Nie,Jiale Cao
出处
期刊:Proceedings of the ... AAAI Conference on Artificial Intelligence [Association for the Advancement of Artificial Intelligence]
卷期号:39 (3): 2735-2743
标识
DOI:10.1609/aaai.v39i3.32278
摘要

Multimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial information between features extracted from 2D images and those derived from 3D point clouds. Existing methods usually aggregate multimodal features at a single stage. However, leveraging multi-stage cross-modal features is crucial for detecting objects of various scales. Therefore, these methods often struggle to integrate features across different scales and modalities effectively, thereby restricting the accuracy of detection. Additionally, the time-consuming Query-Key-Value-based (QKV-based) cross-attention operations often utilized in existing methods aid in reasoning the location and existence of objects by capturing non-local contexts. However, this approach tends to increase computational complexity. To address these challenges, we present SSLFusion, a novel Scale & Space Aligned Latent Fusion Model, consisting of a scale-aligned fusion strategy (SAF), a 3D-to-2D space alignment module (SAM), and a latent cross-modal fusion module (LFM). SAF mitigates scale misalignment between modalities by aggregating features from both images and point clouds across multiple levels. SAM is designed to reduce the inter-modal gap between features from images and point clouds by incorporating 3D coordinate information into 2D image features. Additionally, LFM captures cross-modal non-local contexts in the latent space without utilizing the QKV-based attention operations, thus mitigating computational complexity. Experiments on the KITTI and DENSE datasets demonstrate that our SSLFusion outperforms state-of-the-art methods. Our approach obtains an absolute gain of 2.15% in 3D AP, compared with the state-of-art method GraphAlign on the moderate level of the KITTI test set.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
思源应助酷酷寄容采纳,获得10
刚刚
Jing发布了新的文献求助10
刚刚
天天发布了新的文献求助10
刚刚
天天发布了新的文献求助10
刚刚
鸡蛋卷发布了新的文献求助10
刚刚
天天发布了新的文献求助10
刚刚
1秒前
1秒前
Dylan发布了新的文献求助30
2秒前
煎饼发布了新的文献求助10
2秒前
顾矜应助JJG采纳,获得10
2秒前
知北发布了新的文献求助20
2秒前
2秒前
曾经发布了新的文献求助10
2秒前
1703115完成签到,获得积分10
2秒前
对手完成签到 ,获得积分10
2秒前
维克多关注了科研通微信公众号
2秒前
3秒前
一二三应助文件撤销了驳回
3秒前
卡拉瓦乔完成签到,获得积分10
3秒前
附魔板砖发布了新的文献求助10
3秒前
3秒前
3秒前
杨晓莉莉完成签到,获得积分10
4秒前
Deposit发布了新的文献求助10
4秒前
万能图书馆应助项目采纳,获得10
4秒前
Wayne完成签到,获得积分10
4秒前
温药专完成签到,获得积分10
5秒前
柴佳强发布了新的文献求助10
6秒前
万能图书馆应助yuxuan采纳,获得10
6秒前
科研通AI6.2应助桓桓桓桓采纳,获得10
6秒前
6秒前
临时演员发布了新的文献求助10
7秒前
共享精神应助11011采纳,获得10
8秒前
松尐发布了新的文献求助10
8秒前
8秒前
李嘉怡发布了新的文献求助20
8秒前
yk完成签到,获得积分10
8秒前
桐桐应助仁爱的山兰采纳,获得10
8秒前
科研小白完成签到,获得积分10
8秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Navigating Normative Orders. Interdisciplinary Perspectives 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
A Case Study on Hotels as Noncongregate Emergency Living Accommodations for Returning Citizens 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7757241
求助须知:如何正确求助?哪些是违规求助? 9303638
关于积分的说明 20275298
捐赠科研通 7340757
什么是DOI,文献DOI怎么找? 3311761
关于科研通互助平台的介绍 2462624
邀请新用户注册赠送积分活动 2325463