计算机科学
最大化
人工智能
分布式计算
移动计算
效用最大化
移动电话技术
生成语法
计算机网络
数据建模
领域(数学)
生成模型
机器学习
智能网
数据挖掘
作者
Lianbo Ma,Jiacheng Ding,Ying Qian,Qiang He,Yuanguo Bi,Qing Li
标识
DOI:10.1109/tsc.2026.3680952
摘要
Generative AI (GenAI) has become a research hotspot for the task of content creation and production, which suffers from the issue of high latency due to cloud transmission. One effective solution is to integrate serverless computing with mobile edge computing (MEC) to build a communication-efficient GenAI system, where serverless functions are executed via containers on edge servers. However, the nonnegligible latency of container deployment and cold starts degrades the quality of GenAI service. This issue becomes even more serious in dynamic MEC with mobile and uncertain users. In this paper, we study the provisioning of latency-sensitive query services in GenAI-enabled serverless MEC through dynamic utility maximization. While GenAI of users deployed in cloud refers to as primary GenAI, we deploy their GenAI replicas based on serverless functions in edge servers to maximize user service satisfaction (i.e., utility function). We first formulate a joint decision problem, i.e.,GenAIReplicaAllocation andPlacement (GRAP) problem, under various resource constraints. For this problem, we propose an approximation solver with a provable approximation ratio. Then, we consider an dynamic GRAP problem with uncertain values of users and stochastic request arrivals, and devise a performance-guaranteed online algorithm for a special case of the problem by assuming only a small subset of edge servers suffers significant utility degradation. Finally, we conduct theoretical analysis and experimentation to validate the effectiveness of the proposed mechanisms. Experimental results demonstrate that the proposed mechanisms consistently outperform baseline methods in both service latency and user satisfaction.
科研通智能强力驱动
Strongly Powered by AbleSci AI