计算机科学
吞吐量
卷积神经网络
并行计算
管道(软件)
架空(工程)
推论
GSM演进的增强数据速率
多核处理器
计算
联营
利用
卷积(计算机科学)
边缘设备
计算机工程
算法
人工智能
人工神经网络
无线
操作系统
程序设计语言
云计算
电信
计算机安全
作者
Siqi Wang,Gayathri Ananthanarayanan,Yifan Zeng,Neeraj Goel,Anuj Pathania,Tulika Mitra
标识
DOI:10.1109/tcad.2019.2944584
摘要
IoT Edge intelligence requires Convolutional Neural Network (CNN) inference\nto take place in the edge devices itself. ARM big.LITTLE architecture is at the\nheart of prevalent commercial edge devices. It comprises of single-ISA\nheterogeneous cores grouped into multiple homogeneous clusters that enable\npower and performance trade-offs. All cores are expected to be simultaneously\nemployed in inference to attain maximal throughput. However, high communication\noverhead involved in parallelization of computations from convolution kernels\nacross clusters is detrimental to throughput. We present an alternative\nframework called Pipe-it that employs pipelined design to split convolutional\nlayers across clusters while limiting parallelization of their respective\nkernels to the assigned cluster. We develop a performance-prediction model that\nutilizes only the convolutional layer descriptors to predict the execution time\nof each layer individually on all permitted core configurations (type and\ncount). Pipe-it then exploits the predictions to create a balanced pipeline\nusing an efficient design space exploration algorithm. Pipe-it on average\nresults in a 39% higher throughput than the highest antecedent throughput.\n
科研通智能强力驱动
Strongly Powered by AbleSci AI