计算机科学
稳健性(进化)
强化学习
分布式计算
过程(计算)
弹道
聚类分析
最优化问题
功能(生物学)
模块化设计
人工神经网络
资源配置
全局优化
轨迹优化
代表(政治)
控制(管理)
高效能源利用
样品(材料)
资源管理(计算)
最优控制
自动化
人工智能
功率(物理)
服务(商务)
机器学习
个性化
互操作性
无线
数学优化
运动学
资源(消歧)
利用
作者
Xiangyu Wu,Changbo Hou,Guojing Meng,Zhichao Zhou,Qin Hui Liu
出处
期刊:Drones
[Multidisciplinary Digital Publishing Institute]
日期:2026-02-06
卷期号:10 (2): 116-116
被引量:1
标识
DOI:10.3390/drones10020116
摘要
Multi-unmanned aerial vehicle (UAV) systems are crucial for establishing resilient communication networks in disaster-stricken areas, but their limited energy and dynamic characteristics pose significant challenges for sustained and reliable service provision. Optimizing resource allocation in this situation is a complex sequential decision-making problem, which is naturally suitable for multi-agent reinforcement learning (MARL). However, the most advanced MARL methods (e.g., multi-agent proximal policy optimization (MAPPO)) often encounter difficulties in the “loosely coupled” multi-UAV environment due to their overly centralized evaluation mechanism, resulting in unclear credit assignment and inhibiting personalized optimization. To overcome this, we propose a novel hierarchical framework supported by MAPPO with decoupled critics (MAPPO-DC). Our framework employs an efficient clustering algorithm for user association in the upper layer, while MAPPO-DC is used in the lower layer to enable each UAV to learn customized trajectories and power control strategies. MAPPO-DC achieves a complex balance between global coordination and personalized exploration by redesigning the update rules of the critic network, allowing for precise and personalized credit assignment in a loosely coupled environment. In addition, we designed a composite reward function to guide the learning process towards the goal of proportional fairness. The simulation results show that our proposed MAPPO-DC outperforms existing baselines, including independent proximal policy optimization (IPPO) and standard MAPPO, in terms of communication performance and sample efficiency, validating the effectiveness of our tailored MARL architecture for the task. Through model robustness experiments, we have verified that our proposed MAPPO-DC still has certain advantages in strongly coupled environments.
科研通智能强力驱动
Strongly Powered by AbleSci AI