强化学习
一般化
最大值和最小值
人工智能
背景(考古学)
集合(抽象数据类型)
计算机科学
航程(航空)
功能(生物学)
工作区
机器人
机器学习
数学
工程类
程序设计语言
古生物学
航空航天工程
数学分析
生物
进化生物学
作者
Victor R. F. Miranda,Armando Alves Neto,Gustavo Freitas,Leonardo Amaral Mozelli
标识
DOI:10.1109/tie.2023.3290244
摘要
This paper addresses the application of Deep Reinforcement Learning (DRL) methods in the context of local navigation, i.e., a robot moves towards a goal location in unknown and cluttered workspaces equipped only with limited-range exteroceptive sensors. Collision avoidance policies based on DRL present advantages, but they are quite susceptible to local minima, once their capacity to learn suitable actions is limited to the sensor range. We address this issue by means of reward shaping in actorcritic networks. A dense reward function, that incorporates map information gained in the training stage, is proposed to increase the agent's capacity to decide about the best action. Also, we offer a comparison between the Twin Delayed Deep-Deterministic Policy Gradient (TD3) andSoft Actor-Critic (SAC) algorithms for training our policy. A set of sim-to-sim and sim-to-real trials illustrate that our proposed reward shaping outperforms the compared methods in terms of generalization, by arriving at the target at higher rates in maps that are prone to local minima and collisions.
科研通智能强力驱动
Strongly Powered by AbleSci AI