计算机科学
推论
分布式计算
边缘计算
软件部署
延迟(音频)
移动边缘计算
分拆(数论)
最优化问题
资源配置
移动设备
移动计算
异构网络
云计算
计算资源
边缘设备
计算复杂性理论
资源(消歧)
GSM演进的增强数据速率
无线接入网
移动电话技术
负载平衡(电力)
无线网络
资源管理(计算)
无线
计算机网络
作者
Tong Zheng,Yuanguo Bi,Guangjie Han,Tianao Xiang,Lexi Xu,Qiang He,Liang Zhao
标识
DOI:10.1109/tmc.2026.3650838
摘要
Large language models (LLMs) are revolutionizing various fields due to their powerful generation capabilities. However, their immense computational complexity poses significant challenges in resource consumption, inference latency, and data privacy for traditional cloud-centric deployments. Edge artificial intelligence (Edge-AI) offers promising LLMs deployment solutions by leveraging distributed resources at the network edge. However, existing approaches struggle to adapt to dynamic workloads and efficiently utilize heterogeneous resources in Mobile Edge Computing (MEC) environments. This paper proposes a Dynamic Batching and Adaptive Partitioning (DyBAP) scheme for LLMs deployment, which utilizes ubiquitous geo-distributed resources via end-edge-cloud collaboration. Firstly, we formulate a collaboration deployment optimization problem to minimize inference latency and resource usage under heterogeneous resource and user requirements for latency and accuracy constraints, which is NP-hard. Secondly, to solve this, we develop a dynamic batch fusion optimization algorithm that optimizes the batch size of inference by utilizing the parallel processing power of computing units to balance the latency and resource usage. A block-aware partition optimization algorithm based on multi-agent reinforcement learning (MARL) is proposed for efficient transformer block allocation, integrating mobility awareness for optimal partitioning across dynamic network environments. Simulation results demonstrate the superiority of DyBAP over other benchmarks, reducing inference latency by 17.94% and saving 11.12% in memory resource consumption compared to the end-edge-cloud collaboration approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI