强化学习
一致性(知识库)
计算机科学
理论(学习稳定性)
不稳定性
采样(信号处理)
分解
国家(计算机科学)
控制理论(社会学)
透视图(图形)
价值(数学)
动作(物理)
瞬态(计算机编程)
数学优化
人工智能
机器学习
数学
控制(管理)
算法
计算机视觉
生态学
物理
滤波器(信号处理)
生物
操作系统
量子力学
机械
作者
Haichuan Gao,Zhile Yang,Tian Tan,Tianren Zhang,Jinsheng Ren,Pengfei Sun,Shangqi Guo,Feng Chen
标识
DOI:10.1109/tnnls.2022.3165941
摘要
Undiscounted return is an important setup in reinforcement learning (RL) and characterizes many real-world problems. However, optimizing an undiscounted return often causes training instability. The causes of this instability problem have not been analyzed in-depth by existing studies. In this article, this problem is analyzed from the perspective of value estimation. The analysis result indicates that the instability originates from transient traps that are caused by inconsistently selected actions. However, selecting one consistent action in the same state limits exploration. For balancing exploration effectiveness and training stability, a novel sampling method called last-visit sampling (LVS) is proposed to ensure that a part of actions is selected consistently in the same state. The LVS method decomposes the state-action value into two parts, i.e., the last-visit (LV) value and the revisit value. The decomposition ensures that the LV value is determined by consistently selected actions. We prove that the LVS method can eliminate transient traps while preserving optimality. Also, we empirically show that the method can stabilize the training processes of five typical tasks, including vision-based navigation and manipulation tasks.
科研通智能强力驱动
Strongly Powered by AbleSci AI