阶段(地层学)
计算机科学
计算机视觉
人工智能
目标检测
对象(语法)
模式识别(心理学)
地质学
古生物学
作者
Ahmed Mostayed,Xuefu Zhou
标识
DOI:10.1109/intcec61833.2024.10602914
摘要
In this research, we present a single-shot object detection model designed for far-infrared (FIR) thermal images. Prior research in this field treated thermal modality merely as a supplement to the visible (RGB) images, with a focus on fusing the deep features of the thermal and visible images to improve object detection performance, particularly in low-illumination conditions. This approach, often relying on large, dual stream neural networks for feature extraction leads to impractical model sizes for real-time applications like advanced driver assistance systems (ADAS). It can be argued that prior work overlooked optimizing the training of the thermal imaging stream to achieve the best possible results and relied heavily on the fusion techniques. In this work, we investigate various state-of-the-art techniques for object detection, including efficient architecture design, improved loss functions, prior training on urban scenes, and enhanced post-processing methods. Our best-performing model is based on the YOLO (You Only Look Once) architecture. Trained on the thermal images from the FLIR ADAS dataset, it achieves a remarkable mean average precision (mAP) of 78%, surpassing the previous state-of-the-art by 4%, despite having only two-thirds of the model complexity. Notably, our single-modal model outperforms a previously reported multimodal detection model (combining thermal and RGB inputs) for the same dataset by an impressive margin of 17%.
科研通智能强力驱动
Strongly Powered by AbleSci AI