强化学习
钢筋
逃避(道德)
计算机科学
追逃
人工智能
心理学
社会心理学
生物
免疫系统
免疫学
作者
Xiaoxiao Wang,Peng Yi,Yiguang Hong
标识
DOI:10.1109/tai.2025.3566069
摘要
This paper addresses the problem of Collective Pursuit-Evasion Game (C-PEG) with partial observation, aiming to enhance cooperation among cluster pursuers, reduce the game space, and optimize the overall pursuer strategy efficiency in complex dynamic environments. Firstly, we consider a distributed communication model suitable for cluster collaboration and model the multi-agent cooperative task as a Collective Decentralized Partially Observable Markov Decision Process (C-Dec POMDP). Then we propose a Hierarchical Regularization Game Multi-agent Deep Deterministic Policy Gradient (HRG-MADDPG) algorithm, which includes High-Level Strategy (HLS) and Low-Level Strategy (LLS). The HLS takes the current observation from the pursuers’ sensors and the pursuers’ status as inputs, and outputs target allocation strategies for the coalitions and individuals to guide the LLS in executing specific pursuit tasks. The LLS is further divided into centralized training and decentralized execution phases. In the centralized training phase, the guiding strategy provided by the HLS and the introduced regularization auxiliary term are combined to promote decision-making through the fusion of favorable situations and auxiliary terms. In the decentralized execution phase, a reliable policy improvement method based on partially observable information is designed. All in all, the HRG-MADDPG algorithm provides a reliable method for collective pursuit-evasion. It is trained in a single environment and further validated in multiple different test environments. Simulation results confirm the effectiveness of the proposed method, showing the adaptability, scalability, and rapid responsiveness of the proposed framework in various practical applications.
科研通智能强力驱动
Strongly Powered by AbleSci AI