判别式
计算机科学
多光谱图像
人工智能
模态(人机交互)
模式
模式识别(心理学)
计算机视觉
数据挖掘
机器学习
社会科学
社会学
作者
Shuai You,Xuedong Xie,Yujian Feng,Chaojun Mei,Yimu Ji
标识
DOI:10.1109/lsp.2023.3309578
摘要
Multispectral object detection for autonomous driving is multi-object localization and classification task on visible and thermal modalities. In this scenario, modality differences lead to the lack of object information in a single modality and the misalignment of cross-modality information. To alleviate these problems, most existing methods extract information based on a single scale ( e.g ., these methods mainly focus on detecting significant cars or pedestrians), which leads to insufficient performance in capturing multi-scale discriminative information ( e.g ., small bicycles and blurred pedestrians) and safety hazards in the driving process. In this paper, we propose a Multi-Scale Aggregation Network (MSANet) consisting of two parts Multi-Scale Aggregation Transformer (MSAT) and the Cross-modal Merging Fusion Mechanism (CMFM), which combined with the advantages of Transformer and CNN to extract rich image information from two modalities by mining both local and global context dependencies. Firstly, to reduce the lack of information in a single modality, we design a novel MSAT module to extract rich details and texture from multi-scale. Secondly, to alleviate feature misalignment caused by modality differences, the CMFM is utilized to aggregate complementary information on multiple levels. Comprehensive experiments on two benchmarks demonstrate that our approach shows better results than several state-of-the-art methods. The code is available at https://github.com/ysh-strive/MSANet .
科研通智能强力驱动
Strongly Powered by AbleSci AI