计算机科学
现场可编程门阵列
卷积神经网络
可扩展性
深度学习
片上多核系统
嵌入式系统
人工智能
地标
计算机体系结构
计算机硬件
模式识别(心理学)
芯片上的系统
数据库
作者
Weizhuang Liu,Kejun Tan
标识
DOI:10.1109/icsp54964.2022.9778436
摘要
Convolutional neural network (CNN) has a wide range of applications in face detection and recognition, image classification and semantic segmentation, but it is very difficult to deploy CNN on FPGA embedded platform. The Deep Learning Processor Unit (DPU) released by Xilinx is different from the previous deployment of FPGA, which can accelerate the realization of CNN deployment on FPGA platform and supports a variety of classical CNN structures. In this paper, the face and landmark detection CNN is deployed on ZCU102 platform using DPU based on idea of hardware and software co-design. According to the network features supported by DPU, Normalize network features in VGG-SSD were adjusted to BatchNormalize network features, Convolution was added in LeNet and a double-layer convolution structure was adopted, and the model was pruned to reduce resource consumption and computation. Dual-core DPU and deep flow architecture were used to improve data throughput. The experimental results show that the average detection time of single frame video image face and landmark detection is 26ms, and this design improves the acceleration effect significantly, and has good scalability.
科研通智能强力驱动
Strongly Powered by AbleSci AI