可见的
追逃
逃避(道德)
计算机科学
人工智能
物理
生物
免疫系统
量子力学
免疫学
作者
Yufei Zhuang,Delin Qu,Yiyi Yao,Haibin Huang,R. P. Wang,Lianqi Duan
标识
DOI:10.1109/cac59555.2023.10451200
摘要
The pursuit-evasion game for unmanned surface vessels (USVs) is a topic of great interest in the field of ocean engineering. Most of the previous studies only focused on ideal scenarios with fully observable environments, which are beyond the practical applications. In this paper, we present a reinforcement learning (RL) model for the real marine environment including signal shielding regions, depots, and obstacles. Suppose in the signal shielding regions, the vessels' position information is unknown to each other except for the leader pursuer, and the new environment model is partially observable to ordinary USVs. Then, a distributed multi-agent RL algorithm is proposed and a leader-follower implicit communication mechanism to train the pursuer's strategy is introduced. Weighting the distance, number of captures, and other relevant parameters, a new reward function of the pursuit-evasion game is designed to optimize the pursuit strategy continuously. As shown in the simulation results, the ordinary pursuers can predict the movement direction and capture points of the evader with high accuracy, even when it is in the signal shielding area. This is due to the implicit communication with the leader pursuer in the partially observable environment, which can also effectively optimize the pursuit strategy in the new environment.
科研通智能强力驱动
Strongly Powered by AbleSci AI