计算机科学
推论
移动设备
调度(生产过程)
分布式计算
备品备件
钥匙(锁)
移动计算
推理机
动态优先级调度
移动电话技术
实时计算
业余时间
嵌入式系统
作者
Aoxing Liang,Chengfan Hong,Yunzhe Li,Hongzi Zhu
标识
DOI:10.1109/jiot.2026.3659676
摘要
Large Language Models (LLMs) have demonstrated exceptional natural language understanding and generation capabilities in many emerging applications. However, providing end users with real-time LLM inference on their mobile devices poses substantial challenges. In this work, we propose a fully distributed LLM inference framework, where spare computational resources across heterogeneous mobile devices are effectively harnessed. Our approach integrates two key components: initial model preloading and layer-wise execution scheduling. The proposed methods tackle two major challenges, optimal parameter preloading without prior knowledge of device availability and real-time inference scheduling under dynamic conditions. We evaluate the performance of our framework in real-world environments, experimental results show that it significantly reduces end-to-end delay by 44%∼83% compared to existing approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI