解码方法
计算机科学
现场可编程门阵列
建筑
计算机体系结构
嵌入式系统
多媒体
电信
艺术
视觉艺术
作者
Wenheng Ma,Xinhao Yang,Shulin Zeng,Tengxuan Liu,Libo Shen,Hongyi Wang,Shiyao Li,J.J. Wang,Yuhan Zhang,Hongfu Guo,Jintao Li,Ziming Zhang,Zhenhua Zhu,Xuefei Ning,Tsung-Yi Ho,Guohao Dai,Yu Wang
标识
DOI:10.1145/3706628.3708863
摘要
For large language model (LLM) acceleration, FPGAs face two challenges: insufficient peak computing performance and unacceptable accuracy loss of model compression. This paper proposes FMC-LLM to enable FPGAs for efficient batched decoding of 70B+ LLMs.
科研通智能强力驱动
Strongly Powered by AbleSci AI