强化学习
控制理论(社会学)
计算机科学
理论(学习稳定性)
模型预测控制
控制器(灌溉)
弹道
过程(计算)
人工智能
投影(关系代数)
学习迁移
梯度下降
职位(财务)
自适应控制
跟踪误差
适应(眼睛)
控制(管理)
机器学习
差异(会计)
航程(航空)
操作员(生物学)
控制系统
跟踪(教育)
在线学习
任务(项目管理)
均方误差
钢筋
选择(遗传算法)
控制工程
最小方差无偏估计量
标识
DOI:10.1016/j.trpro.2026.04.077
摘要
Tuning Modle Predictive Control (MPC) weight matrices remains a tedious process that demands considerable expertise. A method is presented that enables a reinforcement learning agent to adjust gains online while enforcing Lyapunov-based bounds guaranteeing that every proposed gain lies inside a provably stable region. The resulting adaptive controller maintains stability regardless of policy network outputs. Across four Unamnnaed Aerial Vehicle (UAV) platforms spanning a 200-fold mass range (27 g to 5.5 kg), tracking improvements of 22–27% are observed on an aggressive 3D fgure-8 trajectory spanning ±4.0m on x and y axes. Position root mean square error (RMSE) is reduced from 0.45–0.55m to 0.33–0.43m across all platforms, with variance reductions of 28–33%. No stability violations occurred throughout 60 evaluation trials. A projection operator clips out-of-bounds gains before they reach the MPC solver, acting as a hard safety layer. Sequential transfer learning reduces per-platform training by 75%. These findings demonstrate that formal stability constraints and learning-based adaptation can coexist effectively.
科研通智能强力驱动
Strongly Powered by AbleSci AI