计算机科学
延迟(音频)
吞吐量
排队
中央处理器
计算机硬件
实时计算
并行计算
操作系统
电信
程序设计语言
无线
作者
Sehyeon Oh,Yongin Kwon,Jemin Lee
出处
期刊:Sensors
[Multidisciplinary Digital Publishing Institute]
日期:2025-02-24
卷期号:25 (5): 1376-1376
被引量:2
摘要
Real-time object detection demands high throughput and low latency, necessitating the use of hardware accelerators. NPU is specialized hardware designed to accelerate the calculation of deep learning models, providing better energy efficiency and parallel processing performance than existing CPUs or GPUs. In particular, it plays an important role in reducing latency and improving processing speed in applications that require real-time processing. In this paper, we construct a real-time object detection system based on YOLOv3, utilizing Neubla’s Antara NPU, and propose two approaches for performance optimization. First, we ensure the continuity of NPU inference by allowing the CPU to process data in advance through double buffering. Second, in a multi-NPU environment, we distribute tasks among NPUs through queue-based processing and analyze the performance limits using Amdahl’s law. Experimental results demonstrate that compared to a CPU-only environment, applying the NPU in single buffering improved throughput by 2.13 times, double buffering by 3.35 times, and in a multi-NPU environment by 4.81 times. Latency decreased by 1.6 times in single and double buffering, and by 1.18 times in the multi-NPU environment. The accuracy remained consistent, with 31.4 mAP on the CPU and 31.8 mAP on the NPU.
科研通智能强力驱动
Strongly Powered by AbleSci AI