强化学习
计算机科学
蒙特卡罗方法
马尔可夫决策过程
超参数
控制器(灌溉)
职位(财务)
增强学习
蒙特卡罗树搜索
数学优化
马尔可夫过程
人工智能
控制理论(社会学)
控制(管理)
数学
财务
经济
生物
农学
统计
标识
DOI:10.1109/iceccme57830.2023.10253420
摘要
The work solves the problem of valve stem position control using a tabular, off-policy Monte Carlo reinforcement learning. The work proposes a model of control valve, involving the stem friction. The control problem is seen as a finite Markov decision process. The reinforcement learning algorithm is a slight modification of the Sutton–Barto Monte Carlo policy estimation method. Such method learns value functions and optimal policies from experience in the form of learning episodes. The modification involves shortening of learning episodes, for their more economical exploitation. As the original algorithm, the method can be used to learn optimal behavior directly from interaction with the environment, with no model of the environment's dynamics. It is easy and efficient to focus such a method on a small subset of states. Such method does not bootstrap and therefore it is less harmed by violations of the Markov property. The proposed control method has been verified by a series of computer simulations, where it is compared with a three-state bang-bang controller. The simulations show that the proposed controller performs stable and as desired. The proposed method requires very few tuning hyperparameters, which are easy to tune. The target policy learns in few minutes, when interacting with the valve model. The work provides insights on possibilities and limitations of the Monte Carlo learning method, which is still unsettled.
科研通智能强力驱动
Strongly Powered by AbleSci AI