人工智能
计算机科学
特征(语言学)
目标检测
图像融合
融合
特征提取
模式识别(心理学)
计算机视觉
传感器融合
人工神经网络
特征检测(计算机视觉)
构造(python库)
行人检测
图像(数学)
面子(社会学概念)
模态(人机交互)
对偶(语法数字)
图像处理
人脸检测
模式
网络体系结构
对象(语法)
可视化
作者
Congying Sun,Jing Zhang,Huinan Guo,Wuxia Zhang
标识
DOI:10.23919/ccc64809.2025.11179380
摘要
Target detection technology plays a pivotal role in various fields such as traffic monitoring, intelligent security, and autonomous driving, and the introduction of multimodal fusion technology has significantly enhanced the target detection efficiency in complex scenarios. However, current multimodal fusion target detection models face dual challenges in fusion strategy and network architecture design: on the one hand, the fusion methods are highly complex and the utilization of multimodal information is inadequate; on the other hand, manually designed network architectures often overlook the potential of noncorresponding level feature fusion, while neural architecture search-based techniques are accompanied by high computational costs and complexity. In response to this, we innovatively propose a Dual Image Feature Fusion (DIFF) method based on a dualbranch YOLOv8 to construct a lightweight and efficient model. This method first adaptively fuses multi-scale features within each modality to achieve non-corresponding level fusion, and then separately processes similar and specific features across different modalities to realize cross-modal fusion. Experimental and data analysis results demonstrate that our proposed fusion method achieves excellent performance on both the VEDAI and FLIR datasets, with mean Average Precision ($m A P$) scores of 51% and 40.7%, respectively. Furthermore, ablation experiments and visual analysis further validate the effectiveness of the DIFF module, providing a novel approach for image feature fusion-driven target detection.
科研通智能强力驱动
Strongly Powered by AbleSci AI