计算机科学
强化学习
稳健性(进化)
过程(计算)
无线传感器网络
控制(管理)
钥匙(锁)
分布式计算
基线(sea)
无线
功能(生物学)
智能交通系统
任务(项目管理)
约束(计算机辅助设计)
智能控制
实时计算
任务分析
控制系统
物联网
过程控制
深度学习
最优化问题
时间限制
智能代理
人工智能
资源管理(计算)
智能传感器
作者
Jian Gu,Yin Wang,Wen Ji
标识
DOI:10.1109/jiot.2025.3650490
摘要
The cooperative control of UAVs in a complex environment is the key and fundamental problem in intelligent transportation and IoT communication systems. The cooperative control policy of the UAVs is challenging to determine due to the non-stationarity of the environment caused by the interactions between the UAVs and the environment under IoT communication. The use of deep learning algorithms based on a digital twin(DT) system to optimize the control policy has shown significant benefits in terms of efficiency and flexibility. However, the DT is prone to being inconsistent with real-world applications due to sensor errors and data delays, which lead to incorrect policy, even if the learning process is convergent. To address this issue, this paper proposes a federated deep reinforcement learning solution to the multi-UAV intelligent control problem, with a newly developed constrained optimization framework. To effectively eliminate the negative impact of uncertainties such as sensor errors and data delays on system performance, the safety constraint function is integrated into the objective function for optimization. The model undergoes training through the constrained distributed federated reinforcement learning framework, avoiding potentially dangerous behavior and thereby improving the safety and robustness of the resulting control strategies. Experimental results have shown that the proposed solution ensures the safety of the strategy, effectively avoids potential dangerous behaviors, and improves the success rate of multi-UAV cooperative control tasks as well as the average task completion time under IoT communication. Experimental results demonstrate that the proposed solution achieves substantial performance improvements while ensuring operational safety. Compared with baseline methods in IoT environments, the algorithm exhibits min 26.7% faster convergence, min 1.8% higher success rate across swarm sizes, and min 12.3% reduction in task completion time.
科研通智能强力驱动
Strongly Powered by AbleSci AI