计算机科学
重新使用
功能(生物学)
价值(数学)
算法
数学优化
数据建模
软件工程
数学
机器学习
生态学
进化生物学
生物
作者
X. J. Chen,洋 大草,Hongfei Yu,Caihua Sun
出处
期刊:IEEE Access
[Institute of Electrical and Electronics Engineers]
日期:2024-01-01
卷期号:12: 120827-120839
被引量:1
标识
DOI:10.1109/access.2024.3450297
摘要
This paper introduces a model, M0RV, that improves the MuZero algorithm through data reuse and loss function optimization. It proposes reusing training trajectories generated by Monte Carlo Tree Search (MCTS) after filtering through an evaluation function trace into the training process, and on this basis, employs the Advantage-Value method to optimize the neural network loss function, ultimately optimizing the training process. A comparative analysis is conducted between the baseline MuZero algorithm, its A0GB algorithm-enhanced variant M0GB, and the further refined M0RV algorithm, across a spectrum of Atari and intricate board games. Notably, M0RV outperforms its predecessors in both the Lunar Lander and Breakout games, as well as in the board game Hex, under consistent steps parameters and unified reward benchmarks. The empirical findings demonstrate that the M0RV model, in comparison to the MuZero model, substantially enhances training efficacy, successfully fulfilling the objective of optimizing the training methodology.
科研通智能强力驱动
Strongly Powered by AbleSci AI