计算机科学
块(置换群论)
卷积神经网络
现场可编程门阵列
修剪
计算机硬件
还原(数学)
帧(网络)
并行计算
嵌入式系统
人工智能
计算机网络
数学
农学
几何学
生物
作者
Hyeong‐Ju Kang,Byung‐Do Yang
出处
期刊:Sensors
[Multidisciplinary Digital Publishing Institute]
日期:2023-09-27
卷期号:23 (19): 8104-8104
摘要
Convolutional neural networks (CNNs) play a crucial role in many EdgeAI and TinyML applications, but their implementation usually requires external memory, which degrades the feasibility of such resource-hungry environments. To solve this problem, this paper proposes memory-reduction methods at the algorithm and architecture level, implementing a reasonable-performance CNN with the on-chip memory of a practical device. At the algorithm level, accelerator-aware pruning is adopted to reduce the weight memory amount. For activation memory reduction, a stream-based line-buffer architecture is proposed. In the proposed architecture, each layer is implemented by a dedicated block, and the layer blocks operate in a pipelined way. Each block has a line buffer to store a few rows of input data instead of a frame buffer to store the whole feature map, reducing intermediate data-storage size. The experimental results show that the object-detection CNNs of MobileNetV1/V2 and an SSDLite variant, widely used in TinyML applications, can be implemented even on a low-end FPGA without external memory.
科研通智能强力驱动
Strongly Powered by AbleSci AI