强化学习
计算机科学
人工智能
自然(考古学)
机器人学
功能(生物学)
差异(会计)
机器学习
机器人
经济
会计
考古
进化生物学
生物
历史
作者
I. Grondman,Lucian Buşoniu,Gabriel A. D. Lopes,Robert Babuška
出处
期刊:IEEE transactions on systems, man and cybernetics
[Institute of Electrical and Electronics Engineers]
日期:2012-11-01
卷期号:42 (6): 1291-1307
被引量:1021
标识
DOI:10.1109/tsmcc.2012.2218595
摘要
Policy-gradient-based actor-critic algorithms are amongst the most popular algorithms in the reinforcement learning framework. Their advantage of being able to search for optimal policies using low-variance gradient estimates has made them useful in several real-life applications, such as robotics, power control, and finance. Although general surveys on reinforcement learning techniques already exist, no survey is specifically dedicated to actor-critic algorithms in particular. This paper, therefore, describes the state of the art of actor-critic algorithms, with a focus on methods that can work in an online setting and use function approximation in order to deal with continuous state and action spaces. After starting with a discussion on the concepts of reinforcement learning and the origins of actor-critic algorithms, this paper describes the workings of the natural gradient, which has made its way into many actor-critic algorithms over the past few years. A review of several standard and natural actor-critic algorithms is given, and the paper concludes with an overview of application areas and a discussion on open issues.
科研通智能强力驱动
Strongly Powered by AbleSci AI