最优控制
数学优化
动态规划
趋同(经济学)
计算机科学
启发式
人工神经网络
强化学习
仿射变换
感知器
状态空间
动作(物理)
控制理论(社会学)
控制(管理)
数学
人工智能
量子力学
统计
物理
经济
经济增长
纯数学
作者
Dongbin Zhao,Zhongpu Xia,Ding Wang
标识
DOI:10.1109/tase.2014.2348991
摘要
In this paper, a self-learning control scheme is proposed for the infinite horizon optimal control of affine nonlinear systems based on the action dependent heuristic dynamic programming algorithm. The policy iteration technique is introduced to derive the optimal control policy with feasibility and convergence analysis. It shows that the "greedy" control action for each state is uniquely existent, the learned control policy after each policy iteration is admissible, and the optimal control policy is able to be obtained. Two three-layer perceptron neural networks are employed to implement the scheme. The critic network is trained by a novel rule to conform to the Bellman equation, and the action network is trained to yield a better control policy. Both training processes alternate until the optimal control policy is achieved. Two simulation examples are provided to validate the effectiveness of the approach. Note to Practitioners - The objective of designing optimal controllers without mathematical models is sought by control practitioners, whereas existing approaches usually derive optimal controllers by accessing the mathematical models or identified models. This paper proposes a new approach which derives optimal controllers by numerical iteration method without accessing any knowledge of the mathematical models. It gives evaluation for every state-action pair in the whole state-action space through the collected data of the underlying system, and then selects the action with the best evaluation for each state. What is required initial admissible control policy. Theorems show that optimal controllers can be acquired and simulation studies verify effectiveness. Further research will extend this approach to online self-learning optimal control approach, thus it can adapt the variation of underlying systems.
科研通智能强力驱动
Strongly Powered by AbleSci AI