PP-PG: Combining Parameter Perturbation with Policy Gradient Methods for Effective and Efficient Explorations in Deep Reinforcement Learning

强化学习 计算机科学 参数统计 摄动(天文学) 动作(物理) 人口 人工智能 数学优化 机器学习 数学 物理 量子力学 统计 社会学 人口学
作者
Shilei Li,Meng Li,Jiongming Su,Shaofei Chen,Zhimin Yuan,Qing Ye
出处
期刊:ACM Transactions on Intelligent Systems and Technology [Association for Computing Machinery]
卷期号:12 (3): 1-21 被引量:1
标识
DOI:10.1145/3452008
摘要

Efficient and stable exploration remains a key challenge for deep reinforcement learning (DRL) operating in high-dimensional action and state spaces. Recently, a more promising approach by combining the exploration in the action space with the exploration in the parameters space has been proposed to get the best of both methods. In this article, we propose a new iterative and close-loop framework by combining the evolutionary algorithm (EA), which does explorations in a gradient-free manner directly in the parameters space with an actor-critic, and the deep deterministic policy gradient (DDPG) reinforcement learning algorithm, which does explorations in a gradient-based manner in the action space to make these two methods cooperate in a more balanced and efficient way. In our framework, the policies represented by the EA population (the parametric perturbation part) can evolve in a guided manner by utilizing the gradient information provided by the DDPG and the policy gradient part (DDPG) is used only as a fine-tuning tool for the best individual in the EA population to improve the sample efficiency. In particular, we propose a criterion to determine the training steps required for the DDPG to ensure that useful gradient information can be generated from the EA generated samples and the DDPG and EA part can work together in a more balanced way during each generation. Furthermore, within the DDPG part, our algorithm can flexibly switch between fine-tuning the same previous RL-Actor and fine-tuning a new one generated by the EA according to different situations to further improve the efficiency. Experiments on a range of challenging continuous control benchmarks demonstrate that our algorithm outperforms related works and offers a satisfactory trade-off between stability and sample efficiency.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
来路遥迢完成签到,获得积分10
刚刚
红绿蓝完成签到 ,获得积分10
刚刚
华仔应助fog采纳,获得10
1秒前
bkagyin应助随风采纳,获得10
1秒前
ER关闭了ER文献求助
1秒前
晚与风i发布了新的文献求助10
1秒前
Jasper应助小刘采纳,获得10
1秒前
hanqianqian发布了新的文献求助10
1秒前
2秒前
silin完成签到,获得积分10
2秒前
深情的依风完成签到,获得积分10
2秒前
微纳组刘同完成签到,获得积分10
2秒前
淮竹完成签到,获得积分10
2秒前
xxy991007完成签到,获得积分10
2秒前
问柳完成签到 ,获得积分10
2秒前
You发布了新的文献求助10
2秒前
唠叨的文龙完成签到,获得积分10
2秒前
3秒前
百里丹珍完成签到,获得积分10
3秒前
3秒前
3秒前
3秒前
波安班完成签到,获得积分10
3秒前
开心人达发布了新的文献求助10
3秒前
酷波er应助waubycid采纳,获得10
3秒前
皇帝的床帘完成签到,获得积分10
3秒前
弘卿发布了新的文献求助10
4秒前
111完成签到,获得积分10
4秒前
Orange应助雨rain采纳,获得30
4秒前
4秒前
588完成签到,获得积分10
4秒前
jiangnan完成签到,获得积分10
4秒前
4秒前
xxy991007发布了新的文献求助10
5秒前
Sugar完成签到,获得积分10
5秒前
yuanmm完成签到,获得积分10
6秒前
球球发布了新的文献求助10
6秒前
xrkxrk完成签到 ,获得积分0
6秒前
6秒前
梅子黄时雨完成签到,获得积分10
6秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Navigating Normative Orders. Interdisciplinary Perspectives 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7759872
求助须知:如何正确求助?哪些是违规求助? 9305126
关于积分的说明 20285682
捐赠科研通 7343898
什么是DOI,文献DOI怎么找? 3312690
关于科研通互助平台的介绍 2463217
邀请新用户注册赠送积分活动 2326657