计算机科学
最大化
计算
分布式计算
数学优化
实时计算
算法
数学
作者
S.-Q. Li,W. Li,Huaguang Shi,W. Y. Yan,Yi Zhou
标识
DOI:10.1145/3640912.3640991
摘要
Unmanned Aerial Vehicles (UAVs) equipped with Multi-Access Edge Computing (MEC) servers can assist Terminal Devices (TDs) in offloading data tasks. In this paper, we investigate a resource allocation and trajectory optimization problem of multiple UAVs assisting TDs in task computation, and our main goal is to improve the task computation efficiency of the system to meet the high-quality experience of TDs. We consider the fairness of TD's computing data volume and the fairness of UAV energy consumption. The problem is transformed into a Partially Observable Markov Decision Process (POMDP). The large action space generated during the UAV flight and resource allocation decision-making process leads to a policy overfitting problem for Multi-Agent Proximal Policy Optimization (MAPPO) method. Policy overfitting causes the UAV to update the policy gradient in the suboptimization direction, preventing it from exploring better flight trajectories. To meet this challenge, we propose a novel method of policy regularization, NV-MAPPO. Simulation results show that NV-MAPPO has significant advantages in latency and energy consumption.
科研通智能强力驱动
Strongly Powered by AbleSci AI