强化学习
计算机科学
钢筋
人工智能
心理学
社会心理学
作者
Homayoun Honari,Mehran Ghafarian Tamizi,Homayoun Najjaran
标识
DOI:10.1109/icra57147.2024.10611316
摘要
Safe reinforcement learning (Safe RL) refers to a class of techniques that\naim to prevent RL algorithms from violating constraints in the process of\ndecision-making and exploration during trial and error. In this paper, a novel\nmodel-free Safe RL algorithm, formulated based on the multi-objective policy\noptimization framework is introduced where the policy is optimized towards\noptimality and safety, simultaneously. The optimality is achieved by the\nenvironment reward function that is subsequently shaped using a safety critic.\nThe advantage of the Safety Optimized RL (SORL) algorithm compared to the\ntraditional Safe RL algorithms is that it omits the need to constrain the\npolicy search space. This allows SORL to find a natural tradeoff between safety\nand optimality without compromising the performance in terms of either safety\nor optimality due to strict search space constraints. Through our theoretical\nanalysis of SORL, we propose a condition for SORL's converged policy to\nguarantee safety and then use it to introduce an aggressiveness parameter that\nallows for fine-tuning the mentioned tradeoff. The experimental results\nobtained in seven different robotic environments indicate a considerable\nreduction in the number of safety violations along with higher, or competitive,\npolicy returns, in comparison to six different state-of-the-art Safe RL\nmethods. The results demonstrate the significant superiority of the proposed\nSORL algorithm in safety-critical applications.\n
科研通智能强力驱动
Strongly Powered by AbleSci AI