波束赋形
计算机科学
弹道
轨迹优化
电信
天文
物理
作者
Qian Gao,Ruikang Zhong,Hyundong Shin,Yuanwei Liu
标识
DOI:10.1109/jiot.2024.3453195
摘要
A multiple unmanned aerial vehicle (UAV) enabled integrated sensing and communication (ISAC) system is investigated. In contrast to existing UAV-enabled ISAC systems assuming static users or 2-D UAV trajectory, we consider a practical roaming user scenario and a 3-D deployment for UAVs. Then, a joint trajectory and beamforming optimization problem is formulated for maximizing the long-term sum data rate, subject to the transmitting power constraint and ensuring beam pattern gain constraint for sensing target. To address the challenge caused by the dynamic and high dimensionality features, multiagent reinforcement learning (MARL) is employed for this partial observation Markov decision process (POMDP) problem. We proposed a two-step approach for against the dynamic scenario: 1) a K-means-based hierarchical user association algorithm is proposed to renew the user association periodically and 2) a hybrid reward multiagent proximal policy optimization (HR-MAPPO) algorithm is proposed, which decomposes the complex combined reward into a team reward and an individual reward. HR-MAPPO introduces a hyperparameter to control the proportion of team/individual action. Numerical results demonstrate that the proposed HR-MAPPO algorithm can outperform the conventional single-agent and multiagent RL algorithms by maintaining high scores on both the sum data rate and beam pattern gain.
科研通智能强力驱动
Strongly Powered by AbleSci AI