植绒(纹理)
强化学习
避碰
固定翼
计算机科学
可扩展性
翼
碰撞
分布式计算
模拟
人工智能
航空航天工程
工程类
计算机安全
物理
数据库
量子力学
作者
Chao Yan,Chang Wang,Han Zhou,Xiaojia Xiang,Xiangke Wang,Lincheng Shen
标识
DOI:10.1109/tits.2024.3505929
摘要
Flocking with multiple unmanned aerial vehicles (UAVs) offers significant potential for diverse applications due to its enhanced maneuverability, improved efficiency, and increased robustness. Collision avoidance is a critical and challenging issue for distributed flocking control with a UAV fleet, especially in dynamic environments with varying numbers of non-cooperative intruders. However, existing reinforcement learning based methods mainly focus on flocking with collision avoidance tasks with static obstacles and a fixed number of UAVs. In this article, we propose a scalable multi-agent reinforcement learning based method to solve the distributed flocking with collision avoidance problem for a scalable fleet of fixed-wing UAVs in dynamic environments. Specifically, we cast this problem in a decentralized partially observable Markov decision process framework and propose a scalable multi-agent reinforcement learning algorithm called spatial-temporal attention multi-agent actor-critic (STAAC). In this algorithm, we design a spatial-temporal attention based population-invariant network architecture to facilitate the representation learning of dynamic dimensional observations. By integrating the local spatial attention and global temporal attention mechanisms, STAAC is able to adapt to the changes in the scale of UAV fleets and the number of intruders. Finally, we empirically demonstrate the effectiveness, scalability, and adaptability of the proposed approach in numerical simulations and hardware-in-the-loop experiments.
科研通智能强力驱动
Strongly Powered by AbleSci AI