Learn Zero-Constraint-Violation Safe Policy in Model-Free Constrained Reinforcement Learning

强化学习 杠杆(统计) 约束(计算机辅助设计) 数学优化 计算机科学 集合(抽象数据类型) 零(语言学) 功能(生物学) 趋同(经济学) 人工智能 数学 经济 语言学 哲学 几何学 进化生物学 生物 程序设计语言 经济增长
作者
Haitong Ma,Changliu Liu,Shengbo Eben Li,Sifa Zheng,Wenchao Sun,Jianyu Chen
出处
期刊:IEEE transactions on neural networks and learning systems [Institute of Electrical and Electronics Engineers]
卷期号:36 (2): 2327-2341 被引量:15
标识
DOI:10.1109/tnnls.2023.3348422
摘要

We focus on learning the zero-constraint-violation safe policy in model-free reinforcement learning (RL). Existing model-free RL studies mostly use the posterior penalty to penalize dangerous actions, which means they must experience the danger to learn from the danger. Therefore, they cannot learn a zero-violation safe policy even after convergence. To handle this problem, we leverage the safety-oriented energy functions to learn zero-constraint-violation safe policies and propose the safe set actor-critic (SSAC) algorithm. The energy function is designed to increase rapidly for potentially dangerous actions, locating the safe set on the action space. Therefore, we can identify the dangerous actions prior to taking them and achieve zero-constraint violation. Our major contributions are twofold. First, we use the data-driven methods to learn the energy function, which releases the requirement of known dynamics. Second, we formulate a constrained RL problem to solve the zero-violation policies. We prove that our Lagrangian-based constrained RL solutions converge to the constrained optimal zero-violation policies theoretically. The proposed algorithm is evaluated on the complex simulation environments and a hardware-in-loop (HIL) experiment with a real autonomous vehicle controller. Experimental results suggest that the converged policies in all environments achieve zero-constraint violation and comparable performance with model-based baseline.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
冰海战记应助久香采纳,获得10
刚刚
今后应助欢呼傲云采纳,获得10
刚刚
fwstu完成签到,获得积分20
刚刚
1秒前
111A完成签到,获得积分10
1秒前
852应助yhy采纳,获得10
2秒前
2秒前
2秒前
3秒前
东方元语应助雷小仙儿采纳,获得20
3秒前
一尔完成签到,获得积分10
4秒前
在水一方应助Jerryer采纳,获得10
5秒前
cdercder应助Shuhan_Song采纳,获得10
6秒前
Mic应助aayu采纳,获得10
6秒前
xxxpeacey发布了新的文献求助10
8秒前
8秒前
10秒前
海娃完成签到 ,获得积分10
10秒前
12秒前
15秒前
野性的柠檬完成签到,获得积分10
16秒前
17秒前
LewisAcid发布了新的文献求助10
17秒前
Lucas应助金刚大王采纳,获得10
19秒前
adi发布了新的文献求助10
19秒前
Azure完成签到 ,获得积分10
19秒前
mumoon完成签到,获得积分10
21秒前
LewisAcid发布了新的文献求助10
21秒前
cdercder应助愉快的枕头采纳,获得10
22秒前
mumoon发布了新的文献求助10
23秒前
23秒前
24秒前
幸福的松鼠完成签到,获得积分10
25秒前
脑洞疼应助甜美的芮采纳,获得10
26秒前
27秒前
Lyzanilia完成签到 ,获得积分10
28秒前
29秒前
东方元语应助qmy采纳,获得20
30秒前
优秀的方盒完成签到 ,获得积分10
30秒前
bkagyin应助高兴的丝采纳,获得10
31秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
An Introduction to Foreign Language Learning and Teaching 750
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 500
What is the Future of Psychotherapy in Digital Age? Technology, AI Bots, and Psychotherapy after Covid 444
煤炭地下气化渗流燃烧方法的研究 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7631373
求助须知:如何正确求助?哪些是违规求助? 9205783
关于积分的说明 19742944
捐赠科研通 7200710
什么是DOI,文献DOI怎么找? 3274592
关于科研通互助平台的介绍 2436554
邀请新用户注册赠送积分活动 2271192