马尔可夫决策过程
计算机科学
强化学习
基站
波束赋形
通信安全
分布式计算
保密
过程(计算)
通信系统
最优化问题
频道(广播)
电信网络
马尔可夫过程
信道状态信息
节点(物理)
无线
安全通信
计算机网络
实时计算
国家(计算机科学)
干扰(通信)
数学优化
稳健性(进化)
不完美的
人工智能
人工神经网络
安全通道
工程类
无线网络
网络安全
无线传感器网络
作者
Zhuang Cao,Di Wu,Tao Fan,Junfeng Zhang,Feng Shu,Edmond Q. Wu
出处
期刊:IEEE Transactions on Vehicular Technology
[Institute of Electrical and Electronics Engineers]
日期:2025-10-20
卷期号:75 (4): 6300-6312
标识
DOI:10.1109/tvt.2025.3623316
摘要
Unmanned aerial vehicle (UAV) plays an increasingly vital role in communication networks. However, ensuring the security of UAV communications remains a significant challenge. This study investigates a secure communication system that employs a reconfigurable intelligent surface (RIS) to assist a UAV base station (UAV-BS), while accounting for the threats posed by eavesdropper and jammer. To enhance the security of UAV-BS communications under imperfect channel state information, we formulate the system as a Markov decision process and propose a method that jointly optimizes the UAV-BS beamforming matrix, trajectory, and RIS phase shift matrix. The coupling of these optimization variables complicates the solution of the nonconvex optimization problem. To address this, we propose the mean target Q-value hierarchical twin delayed deep deterministic policy gradient (MTQ-HTD3) algorithm. The algorithm decomposes the problem into two subproblems using hierarchical reinforcement learning: one focuses on optimizing the UAV-BS beamforming matrix and RIS phase shift matrix, while the other optimizes the UAV-BS trajectory. Additionally, we employ the mean value of multiple target Q-values in the critic network to minimize estimation bias and enhance network stability. Simulation results demonstrate that the proposed MTQ-HTD3 method significantly improves the secrecy rate of the RIS-assisted UAV-BS secure communication system.
科研通智能强力驱动
Strongly Powered by AbleSci AI