Joint Trajectory and Power Optimization for Loosely Coupled Tasks: A Decoupled-Critic MAPPO Approach

计算机科学 稳健性(进化) 强化学习 分布式计算 过程(计算) 弹道 聚类分析 最优化问题 功能(生物学) 模块化设计 人工神经网络 资源配置 全局优化 轨迹优化 代表(政治) 控制(管理) 高效能源利用 样品(材料) 资源管理(计算) 最优控制 自动化 人工智能 功率(物理) 服务(商务) 机器学习 个性化 互操作性 无线 数学优化 运动学 资源(消歧) 利用
作者
Xiangyu Wu,Changbo Hou,Guojing Meng,Zhichao Zhou,Qin Hui Liu
出处
期刊:Drones [Multidisciplinary Digital Publishing Institute]
卷期号:10 (2): 116-116 被引量:1
标识
DOI:10.3390/drones10020116
摘要

Multi-unmanned aerial vehicle (UAV) systems are crucial for establishing resilient communication networks in disaster-stricken areas, but their limited energy and dynamic characteristics pose significant challenges for sustained and reliable service provision. Optimizing resource allocation in this situation is a complex sequential decision-making problem, which is naturally suitable for multi-agent reinforcement learning (MARL). However, the most advanced MARL methods (e.g., multi-agent proximal policy optimization (MAPPO)) often encounter difficulties in the “loosely coupled” multi-UAV environment due to their overly centralized evaluation mechanism, resulting in unclear credit assignment and inhibiting personalized optimization. To overcome this, we propose a novel hierarchical framework supported by MAPPO with decoupled critics (MAPPO-DC). Our framework employs an efficient clustering algorithm for user association in the upper layer, while MAPPO-DC is used in the lower layer to enable each UAV to learn customized trajectories and power control strategies. MAPPO-DC achieves a complex balance between global coordination and personalized exploration by redesigning the update rules of the critic network, allowing for precise and personalized credit assignment in a loosely coupled environment. In addition, we designed a composite reward function to guide the learning process towards the goal of proportional fairness. The simulation results show that our proposed MAPPO-DC outperforms existing baselines, including independent proximal policy optimization (IPPO) and standard MAPPO, in terms of communication performance and sample efficiency, validating the effectiveness of our tailored MARL architecture for the task. Through model robustness experiments, we have verified that our proposed MAPPO-DC still has certain advantages in strongly coupled environments.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
艾尔伯特发布了新的文献求助10
刚刚
1秒前
文在否完成签到,获得积分10
2秒前
鲤鱼听荷发布了新的文献求助10
2秒前
xxx发布了新的文献求助10
3秒前
芒果十六发布了新的文献求助10
3秒前
4秒前
饱满的秋灵完成签到,获得积分10
4秒前
5秒前
yyyyyy发布了新的文献求助10
5秒前
6秒前
6秒前
7秒前
Tonsil01发布了新的文献求助25
7秒前
李晨光发布了新的文献求助30
8秒前
8秒前
YoRHac发布了新的文献求助10
8秒前
8秒前
深情安青的应助被成长的点滴采纳,获得10
9秒前
9秒前
9秒前
悦耳的映寒完成签到,获得积分10
9秒前
每天都很开心完成签到,获得积分10
10秒前
10秒前
10秒前
干净的琦发布了新的文献求助10
10秒前
Tonsil01发布了新的文献求助20
10秒前
赘婿的应助被fpc采纳,获得10
10秒前
小胖完成签到,获得积分10
11秒前
11秒前
mie发布了新的文献求助10
11秒前
天真的小珍完成签到,获得积分20
12秒前
所所的应助被杨江丽采纳,获得10
12秒前
12秒前
12秒前
13秒前
李存完成签到,获得积分10
13秒前
闻屿完成签到,获得积分10
13秒前
Tonsil01发布了新的文献求助25
13秒前
大模型的应助被Arisophila采纳,获得10
14秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
CODESSA Version 2.13 for Windows 2000
Agricultural Ecology (Liao Yuncheng & Lin Wenxiong) 1000
Rosenblum, Global Change Biology 800
Berberine regulates the TLR4 signaling pathway to suppress hypoxia-induced proliferation and migration of pulmonary arterial smooth muscle cells 520
Organizational Behavior 510
Derham on the Law of Set Off (德勒姆论抵消法/第五版) 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 有机化学 化学工程 内科学 物理 生物化学 复合材料 催化作用 细胞生物学 人工智能 心理学 无机化学 基因 遗传学
热门帖子
关注 科研通微信公众号,转发送积分 7845186
求助须知:如何正确求助?哪些是违规求助? 9365579
关于积分的说明 20646545
捐赠科研通 7441346
什么是DOI,文献DOI怎么找? 3341343
关于科研通互助平台的介绍 2485217
邀请新用户注册赠送积分活动 2363781