钢筋
强化学习
人工智能
心理学
认知科学
计算机科学
社会心理学
出处
期刊:
[Institution of Engineering and Technology]
日期:2025-04-09
卷期号:2025 (2): 130-135
标识
DOI:10.1049/icp.2025.1024
摘要
With the development of unmanned equipment, the research on multi-agent systems has gradually become a hot research area in automation. For the problem of multi-agent cooperative roundup, this paper introduces a method based on Reinforcement Learning (RL) into the scenario and obtains a cooperative roundup strategy that does not require complex mathematical models. Firstly, we model the roundup scenario and design the observation space, action space, and reward function of the roundup environment based on the cooperative criterion. The escaping agent uses the artificial potential field method to escape. This paper proposes a new team reward distribution method based on the Hungarian algorithm to effectively solve the problem of agent falling behind. Subsequently, the Multi Agent Deep Deterministic Policy Gradient algorithm (MADDPG) is improved. To address the problems of low sample utilization and inaccurate Q value, we introduce the prioritized experience replay mechanism and the Dueling network. At the same time, the Gumbel-Softmax sampling strategy is utilized to expand the algorithm to discrete action space. In the simulation experiment, the algorithm shows good convergence and effectiveness. The roundup strategy proposed in this paper can successfully complete the multi-agent cooperative roundup task.
科研通智能强力驱动
Strongly Powered by AbleSci AI