Credit of optimal state transition based reinforcement learning algorithm
作者
Tingfeng Bai,Gengfeng Wu
标识
DOI:10.1109/icnnsp.2003.1279213
摘要
This paper proposed an optimal model based on the distance between current state and goal state and the cost of state transition in order to solve goal state problem more effectively. Based on the optimal model, a unique reinforcement learning algorithm named COSTRLA (credit of optimal state transition based reinforcement learning algorithm) is also presented. The COSTRLA defined a COST function used to evaluate optimality of output strategy, developed update principle for the COST function based on the dynamic programming principle, while reinforcement signal is defined as the distance from current state to goal state. The COSTRLA was applied into cooperative control of Buddy-Arnolds robot. The simulation experiment has shown the advantages of COSTRLA over some popular reinforcement learning algorithms such as Q-learning and prioritized sweeping algorithms.