计算机科学
建筑
卷积神经网络
序列(生物学)
对象(语法)
人工智能
状态空间
网络体系结构
国家(计算机科学)
实时计算
模式识别(心理学)
算法
计算机网络
艺术
视觉艺术
统计
生物
遗传学
数学
标识
DOI:10.1016/j.engappai.2025.111067
摘要
Real-time performance is essential for practical deployment of object detection on edge devices, where high processing speed and low latency are paramount. This paper introduces a novel approach aimed at boosting real-time object detection while strictly adhering to computational constraints. A structured state space sequence model, Mamba, is strategically embedded in the early stages of the backbone network to capture long-range dependencies, thereby enhancing the model’s representation capability. Given the limitations of Mamba in directional perception, a lightweight spatial attention mechanism is introduced to integrate global context into each spatial location. Additionally, a computationally efficient module inspired by the Ghost module is developed to reduce resource demands. This dual-strategy approach optimizes both performance and efficiency in real-time object detection. Extensive experiments demonstrate the superiority of this proposed approach; on the Microsoft Common Objects in Context (MS COCO) dataset, it achieves a +1.6 AP (Average Precision) improvement over state-of-the-art methods, reaching 41.1 AP with minimal added model complexity on the nano scale. The effectiveness and efficiency of each component are further substantiated through ablation studies on the Pascal Visual Object Classes (Pascal VOC dataset). To verify the universality of the proposed method, this study selects underwater object detection, characterized by an extremely complex background environment, as the other validation scenario. Through the application of this proposed approach to underwater object detection, a state-of-the-art result of 69.5 AP was obtained on the Detecting Underwater Objects (DUO) dataset, exceeding that of You Only Look Once Detector version 11 (YOLO11) by +0.3 AP. Code: https://github.com/chenjie04/Hybrid-YOLO . • Proposes a novel lightweight spatial attention module is to extract global context. • Proposes a Mamba module complemented with global context adapted for vision tasks. • Introduces cost-effective Ghost Module to reduce parameters while maintaining performance • Presents hybrid Mamba-CNN architecture for real-time object detection enhancement.
科研通智能强力驱动
Strongly Powered by AbleSci AI