计算机科学
强化学习
透视图(图形)
障碍物
集合(抽象数据类型)
多样性(控制论)
简单(哲学)
控制(管理)
分布式计算
国家(计算机科学)
人工智能
机器学习
算法
认识论
哲学
程序设计语言
法学
政治学
作者
Gabriel Barth-Maron,Matthew W. Hoffman,David Budden,Will Dabney,Dan Horgan,Dhruva Tb,Alistair Muldal,Nicolas Heess,Timothy Lillicrap
标识
DOI:10.48550/arxiv.1804.08617
摘要
This work adopts the very successful distributional perspective on reinforcement learning and adapts it to the continuous control setting. We combine this within a distributed framework for off-policy learning in order to develop what we call the Distributed Distributional Deep Deterministic Policy Gradient algorithm, D4PG. We also combine this technique with a number of additional, simple improvements such as the use of $N$-step returns and prioritized experience replay. Experimentally we examine the contribution of each of these individual components, and show how they interact, as well as their combined contributions. Our results show that across a wide variety of simple control tasks, difficult manipulation tasks, and a set of hard obstacle-based locomotion tasks the D4PG algorithm achieves state of the art performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI