非线性系统
计算机科学
强化学习
仿射变换
随机逼近
算法
迭代法
数学优化
人工神经网络
最优控制
数学
人工智能
物理
量子力学
纯数学
计算机安全
钥匙(锁)
作者
Li Wang,Xiushan Jiang,Dongya Zhao,Weihai Zhang
标识
DOI:10.1109/ntci60157.2023.10403711
摘要
In this paper, a synchronous iterative algorithm based on reinforcement learning is proposed for solving the optimal control problem of stochastic affine nonlinear systems subject to state-dependent noise. The algorithm solves the related Hamilton-Jacobi equations online, which are nonlinear and difficult to solve directly, especially when higher-order nonlinear systems are involved. The algorithm is constructed based on an actor/critic structure, which permits the critic and actor neural networks (NNs) formed by value function approximation to be updated synchronously. Persistence of excitation is applied to stochastic nonlinear systems. The author presents an implementation of the algorithm and proves that the NN weights converge to the ideal values. A simulation example validates the effectiveness of the proposed algorithm.
科研通智能强力驱动
Strongly Powered by AbleSci AI