PP-PG: Combining Parameter Perturbation with Policy Gradient Methods for Effective and Efficient Explorations in Deep Reinforcement Learning

强化学习 计算机科学 参数统计 摄动(天文学) 动作(物理) 人口 人工智能 数学优化 机器学习 数学 物理 量子力学 统计 社会学 人口学
作者
Shilei Li,Meng Li,Jiongming Su,Shaofei Chen,Zhimin Yuan,Qing Ye
出处
期刊:ACM Transactions on Intelligent Systems and Technology [Association for Computing Machinery]
卷期号:12 (3): 1-21 被引量:1
标识
DOI:10.1145/3452008
摘要

Efficient and stable exploration remains a key challenge for deep reinforcement learning (DRL) operating in high-dimensional action and state spaces. Recently, a more promising approach by combining the exploration in the action space with the exploration in the parameters space has been proposed to get the best of both methods. In this article, we propose a new iterative and close-loop framework by combining the evolutionary algorithm (EA), which does explorations in a gradient-free manner directly in the parameters space with an actor-critic, and the deep deterministic policy gradient (DDPG) reinforcement learning algorithm, which does explorations in a gradient-based manner in the action space to make these two methods cooperate in a more balanced and efficient way. In our framework, the policies represented by the EA population (the parametric perturbation part) can evolve in a guided manner by utilizing the gradient information provided by the DDPG and the policy gradient part (DDPG) is used only as a fine-tuning tool for the best individual in the EA population to improve the sample efficiency. In particular, we propose a criterion to determine the training steps required for the DDPG to ensure that useful gradient information can be generated from the EA generated samples and the DDPG and EA part can work together in a more balanced way during each generation. Furthermore, within the DDPG part, our algorithm can flexibly switch between fine-tuning the same previous RL-Actor and fine-tuning a new one generated by the EA according to different situations to further improve the efficiency. Experiments on a range of challenging continuous control benchmarks demonstrate that our algorithm outperforms related works and offers a satisfactory trade-off between stability and sample efficiency.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
凛凛发布了新的文献求助10
刚刚
阴暗蘑菇发布了新的文献求助10
2秒前
科研通AI6.4应助司梦黎采纳,获得10
2秒前
花海发布了新的文献求助10
3秒前
朱莉发布了新的文献求助10
3秒前
旺旺仙贝发布了新的文献求助10
4秒前
凉雨渲完成签到,获得积分10
4秒前
horsam完成签到,获得积分10
4秒前
十九集完成签到 ,获得积分10
4秒前
5秒前
洛书发布了新的文献求助10
5秒前
Rosaline完成签到,获得积分10
5秒前
兜里有糖完成签到,获得积分10
5秒前
5秒前
MingM发布了新的文献求助10
6秒前
丰富的寒蕾完成签到 ,获得积分10
7秒前
HUIMAO003完成签到,获得积分10
7秒前
Emma完成签到,获得积分10
7秒前
cross_dream完成签到 ,获得积分10
8秒前
路人发布了新的文献求助10
9秒前
DJ发布了新的文献求助10
10秒前
隐形寒香完成签到 ,获得积分10
10秒前
IWJL发布了新的文献求助10
10秒前
老牛爱耕田完成签到 ,获得积分10
10秒前
dihou111完成签到,获得积分10
10秒前
fx应助Hananx采纳,获得20
11秒前
书记完成签到,获得积分10
12秒前
CipherSage应助无名666采纳,获得10
13秒前
司梦黎完成签到,获得积分20
14秒前
14秒前
15秒前
万能图书馆应助IWJL采纳,获得10
15秒前
小蘑菇应助jzzj采纳,获得10
15秒前
陈小桥完成签到,获得积分10
16秒前
16秒前
hoosen2008完成签到,获得积分20
17秒前
zhuzhu发布了新的文献求助10
19秒前
邓宸峰发布了新的文献求助10
20秒前
核桃应助菠菜采纳,获得30
20秒前
科研蛀虫完成签到,获得积分10
20秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Navigating Normative Orders. Interdisciplinary Perspectives 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7761400
求助须知:如何正确求助?哪些是违规求助? 9306418
关于积分的说明 20294439
捐赠科研通 7345946
什么是DOI,文献DOI怎么找? 3313135
关于科研通互助平台的介绍 2463437
邀请新用户注册赠送积分活动 2327377