强化学习
鞅(概率论)
计算机科学
等价(形式语言)
数学优化
样品(材料)
人工智能
数学
应用数学
色谱法
离散数学
化学
作者
Yang Gu,Yuhu Cheng,Kun Yu,Xuesong Wang
标识
DOI:10.1109/tcyb.2022.3170355
摘要
Since the sample data after one exploration process can only be used to update network parameters once in on-policy deep reinforcement learning (DRL), a high sample efficiency is necessary to accelerate the training process of on-policy DRL. In the proposed method, a submartingale criterion is proposed on the basis of the equivalence relationship between the optimal policy and martingale, and then an advanced value iteration (AVI) method is proposed to conduct value iteration with a high accuracy. Based on this foundation, an anti-martingale (AM) reinforcement learning framework is established to efficiently select the sample data that is conducive to policy optimization. In succession, an AM proximal policy optimization (AMPPO) method, which combines the AM framework with proximal policy optimization (PPO), is proposed to reasonably accelerate the updating process of state value that satisfies the submartingale criterion. Experimental results on the Mujoco platform show that AMPPO can achieve better performance than several state-of-the-art comparative DRL methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI