强化学习
计算机科学
机器学习
人工智能
动作选择
动作(物理)
选择偏差
光学(聚焦)
统计
数学
物理
量子力学
神经科学
光学
感知
生物
作者
Junwei Zhang,Shuai Han,Xi Xiong,Sheng Zhu,Shuai Lü
标识
DOI:10.1016/j.ins.2024.120255
摘要
Actor-critic deep reinforcement learning methods have demonstrated significant performance in many challenging decision-making and control tasks, but also suffer from high sample complexity and overestimation bias. Current researches focus on using underestimation to balance overestimation and reducing bias through ensemble learning, but introducing underestimation bias and excessive network costs. In this paper, we first analyze the effect of action selection policy on estimation bias. Then, we propose the Explorer-Actor-Critic (EAC) method that gives a more conservative objective for the actor to reduce overestimation, introduces a learnable explorer to improve exploration ability, and uses an action mixing mechanism to mitigate experience distribution bias. Furthermore, we apply the EAC method to TD3 and SAC and verify its effectiveness through extensive comparison and ablation experiments. Our algorithm not only outperforms state-of-the-art algorithms, but also is compatible with other actor-critic methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI