微分博弈
马尔可夫决策过程
强化学习
计算机科学
集合(抽象数据类型)
差速器(机械装置)
过程(计算)
功能(生物学)
人工智能
马尔可夫过程
数学优化
工程类
数学
操作系统
程序设计语言
航空航天工程
统计
生物
进化生物学
作者
Jacob T. English,Jay Wilhelm
摘要
Deep reinforcement learning was used to train an agent within the framework of a Markov decision process (MDP) to pursue a target, while avoiding a defender, for the target–attacker–defender (TAD) differential game of pursuit and evasion. The aim of this work was to explore the games where the previous attacking guidance methods found in literature failed to capture the target. The reward function of the MDP presented by this work allowed for an attacking agent to learn a policy that expanded the number of cases where the target is captured beyond the former limit of success through the application of the twin delayed deep deterministic policy gradient algorithm. The strategy developed using artificial intelligence expands the target capture guidance approach to enable the attacker to avoid the defender in states where the two agents are in close proximity. Initial target positions within a limited set were considered with fixed values for agent velocities and attacker and defender initial positions to evaluate the attacker’s learned behavior in comparison with the optimal point capture guidance laws for target capture in the TAD game.
科研通智能强力驱动
Strongly Powered by AbleSci AI