趋同(经济学)
计算机科学
国家(计算机科学)
极小极大
理论(学习稳定性)
跟踪(教育)
控制理论(社会学)
数学优化
强化学习
控制(管理)
最优控制
期限(时间)
Riccati方程
障碍物
线性系统
人工智能
数学
控制器(灌溉)
控制系统
弹道
状态估计器
算法
动作(物理)
跟踪系统
作者
Mingxiang Liu,Qianqian Cai,Wei Meng,Dandan Li,Minyue Fu
标识
DOI:10.1109/tcyb.2026.3652143
摘要
This article investigates the finite-horizon $H_{\infty } $ tracking control problem for discrete-time (DT) linear systems with partial observation and unknown dynamics from a game-theoretic perspective. Unlike existing reinforcement learning (RL) approaches that primarily address infinite-horizon, time-invariant systems with full state information, our setting requires solving time-varying Riccati equations and developing model-free methods that rely solely on input-output data. To tackle these challenges, we reconstruct the system state from historical input-output trajectories, driving to a data-driven system representation, and we define an input-output-based time-varying $Q$ -function. We then propose two minimax $Q$ -learning algorithms that do not require an initially admissible policy and avoid the use of a discount factor, thereby removing a long-standing obstacle to stability guarantees. Moreover, the framework readily extends to both infinite-horizon and time-varying systems without structural modifications. Convergence is proved theoretically, and the effectiveness of the algorithms is validated through simulations.
科研通智能强力驱动
Strongly Powered by AbleSci AI