计算机科学
现场可编程门阵列
数据流
带宽(计算)
卷积神经网络
内存带宽
炸薯条
计算机硬件
建筑
嵌入式系统
并行计算
计算机体系结构
人工智能
计算机网络
艺术
视觉艺术
电信
作者
Yan Chen,Kiyofumi Tanaka
标识
DOI:10.1145/3626202.3637582
摘要
In recent years, Convolutional Neural Networks (CNNs) have been widely used in many fields. Many accelerators are designed to run CNN inference on hardware. They are classified into two types: Overlay and Dataflow. In this paper, we propose a new architecture of CNN Inference accelerators to leverage the strengths of each type. It is flexible, fast, and has low off-chip memory bandwidth requirements. The accelerator based on our new architecture can be configured to fit the CNN structure and FPGA capacity. It basically runs CNNs by blocks, whereas it has a fallback path to run single Convolution or Fully Connected layers efficiently. Our architecture demonstrates significant advantages when running networks composed of inverted residual blocks. We archived 1954FPS (frames per second) when running 8-bit quantized MobileNetV2 on a mid-range FPGA ZU7EV, and the off-chip memory bandwidth requirement is as low as 2.47GiB/s. It also runs on cost-optimized FPGAs, and archived 505FPS on ZU3EG. The throughputs of ours are higher than the other types, while the minimum FPGA capacity and the off-chip memory bandwidth requirement of ours are reasonable.
科研通智能强力驱动
Strongly Powered by AbleSci AI