亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

PP-PG: Combining Parameter Perturbation with Policy Gradient Methods for Effective and Efficient Explorations in Deep Reinforcement Learning

强化学习 计算机科学 参数统计 摄动(天文学) 动作(物理) 人口 人工智能 数学优化 机器学习 数学 物理 量子力学 统计 社会学 人口学
作者
Shilei Li,Meng Li,Jiongming Su,Shaofei Chen,Zhimin Yuan,Qing Ye
出处
期刊:ACM Transactions on Intelligent Systems and Technology [Association for Computing Machinery]
卷期号:12 (3): 1-21 被引量:1
标识
DOI:10.1145/3452008
摘要

Efficient and stable exploration remains a key challenge for deep reinforcement learning (DRL) operating in high-dimensional action and state spaces. Recently, a more promising approach by combining the exploration in the action space with the exploration in the parameters space has been proposed to get the best of both methods. In this article, we propose a new iterative and close-loop framework by combining the evolutionary algorithm (EA), which does explorations in a gradient-free manner directly in the parameters space with an actor-critic, and the deep deterministic policy gradient (DDPG) reinforcement learning algorithm, which does explorations in a gradient-based manner in the action space to make these two methods cooperate in a more balanced and efficient way. In our framework, the policies represented by the EA population (the parametric perturbation part) can evolve in a guided manner by utilizing the gradient information provided by the DDPG and the policy gradient part (DDPG) is used only as a fine-tuning tool for the best individual in the EA population to improve the sample efficiency. In particular, we propose a criterion to determine the training steps required for the DDPG to ensure that useful gradient information can be generated from the EA generated samples and the DDPG and EA part can work together in a more balanced way during each generation. Furthermore, within the DDPG part, our algorithm can flexibly switch between fine-tuning the same previous RL-Actor and fine-tuning a new one generated by the EA according to different situations to further improve the efficiency. Experiments on a range of challenging continuous control benchmarks demonstrate that our algorithm outperforms related works and offers a satisfactory trade-off between stability and sample efficiency.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
Kao应助科研通管家采纳,获得10
47秒前
Kao应助科研通管家采纳,获得10
47秒前
Kao应助科研通管家采纳,获得10
47秒前
Kao应助科研通管家采纳,获得10
47秒前
龙箫羽笛完成签到 ,获得积分10
1分钟前
机灵的以筠完成签到 ,获得积分10
1分钟前
FeelingUnreal完成签到,获得积分10
1分钟前
GHOSTagw完成签到,获得积分10
1分钟前
魔术师完成签到,获得积分10
2分钟前
开放的乐驹完成签到 ,获得积分10
2分钟前
Kao应助科研通管家采纳,获得10
2分钟前
2分钟前
2分钟前
顾矜应助可靠小甜瓜采纳,获得10
3分钟前
Criminology34完成签到,获得积分0
3分钟前
Ttimer完成签到,获得积分10
3分钟前
科研通AI6.2应助雅芳采纳,获得10
4分钟前
4分钟前
雅芳发布了新的文献求助10
4分钟前
Kao应助科研通管家采纳,获得10
4分钟前
5分钟前
LINDENG2004完成签到 ,获得积分10
5分钟前
研友_nEoDm8发布了新的文献求助10
5分钟前
大医仁心完成签到 ,获得积分10
5分钟前
momo完成签到 ,获得积分10
5分钟前
可爱的函函应助研友_nEoDm8采纳,获得10
5分钟前
xinxin完成签到,获得积分10
5分钟前
坦率如之完成签到,获得积分10
5分钟前
自信大树完成签到,获得积分10
6分钟前
风息完成签到,获得积分10
6分钟前
Kao应助科研通管家采纳,获得10
6分钟前
李健应助科研通管家采纳,获得10
6分钟前
Kao应助科研通管家采纳,获得10
6分钟前
无限的盼晴完成签到 ,获得积分10
7分钟前
多情的涔完成签到,获得积分10
7分钟前
drhkc完成签到,获得积分10
7分钟前
专注的小白菜完成签到,获得积分10
8分钟前
8分钟前
8分钟前
Kao应助科研通管家采纳,获得10
8分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Geist der Kunst und Kultur 1000
Resistance Spot Welding Dataset for Automobile Body-in-White Quality Analysis 748
日本現代怪異事典 副読本 700
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 650
Machine Learning for Asset Management and Pricing 600
Numerical analysis of the coupled atmosphere-ocean models (CAO II). II 600
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7392267
求助须知:如何正确求助?哪些是违规求助? 8998427
关于积分的说明 19149455
捐赠科研通 7028247
什么是DOI,文献DOI怎么找? 3229218
关于科研通互助平台的介绍 2391567
邀请新用户注册赠送积分活动 2210581