计算机科学
推论
加速
延迟(音频)
软件部署
隐藏物
灵敏度(控制系统)
门控
国际商用机器公司
软件工程
人工智能
计算机网络
并行计算
工程类
生理学
电信
生物
纳米技术
材料科学
电子工程
作者
Shuzhang Zhong,Ling Liang,Yuan Wang,Runsheng Wang,Ru Huang,Meng Li
标识
DOI:10.1145/3676536.3676741
摘要
Mixture-of-Experts (MoE) models are designed to enhance the efficiency of large language models (LLMs) without proportionally increasing the computational demands. However, their deployment on edge devices still faces significant challenges due to high on-demand loading overheads from managing sparsely activated experts. This paper introduces AdapMoE, an algorithm-system co-design framework for efficient MoE inference. AdapMoE features adaptive expert gating and management to reduce the on-demand loading overheads. We observe the heterogeneity of experts loading across layers and tokens, based on which we propose a sensitivity-based strategy to adjust the number of activated experts dynamically. Meanwhile, we also integrate advanced prefetching and cache management techniques to further reduce the loading latency. Through comprehensive evaluations on various platforms, we demonstrate AdapMoE consistently outperforms existing techniques, reducing the average number of activated experts by 25% and achieving a 1.35x speedup without accuracy degradation. Code is available at: https://github.com/PKU-SEC-Lab/AdapMoE.
科研通智能强力驱动
Strongly Powered by AbleSci AI