计算机科学
人工智能
强化学习
机器人
弹道
机器学习
机器人学
稀缺
钥匙(锁)
夹持器
鲁棒控制
机械臂
马尔可夫决策过程
任务分析
离线学习
控制(管理)
机器人运动学
工作(物理)
稳健性(进化)
在线学习
作者
Botao Dong,Xin Dong,Xi Chen,Mingxuan Wang,Hongtian Chen
标识
DOI:10.1109/tii.2025.3627122
摘要
Offline reinforcement learning (RL) offers a promising paradigm for learning policies from precollected datasets. Nonetheless, applying it to robotic control poses significant challenges, including non-Markovian dynamics and the scarcity of high-quality demonstrations, both of which can undermine the performance of existing methods. To address these issues, this work introduces the horizon-Greedy Q-ensembles Regularized decision Transformer (GQRT), an offline RL algorithm tailored for robotic tasks. GQRT leverages a transformer-based policy for trajectory modeling, thereby enabling effective long-horizon decision-making in non-Markovian settings. To alleviate the lack of expert demonstrations, we develop a multistep horizon-greedy policy evaluation mechanism that stitches suboptimal sequences into improved trajectories. Furthermore, to cope with mixed-quality demonstrations collected from diverse sources, we incorporate a Q-ensemble with lower confidence bound regularization, which ensures more stable and reliable value estimation. Extensive experiments on robotic locomotion and manipulation benchmarks demonstrate that GQRT achieves state-of-the-art performance, validating its robustness in complex robotic scenarios.
科研通智能强力驱动
Strongly Powered by AbleSci AI