计算机科学
机器人
强化学习
过程(计算)
启发式
稳健性(进化)
机器人运动
启发式
控制理论(社会学)
模拟
人工智能
控制(管理)
机器人控制
移动机器人
操作系统
生物化学
化学
基因
作者
Daoling Qin,Guoteng Zhang,Zhengguo Zhu,Teng Chen,Weiliang Zhu,Xuewen Rong,Anhuan Xie,Yibin Li
标识
DOI:10.1142/s0219843623500135
摘要
A new method is proposed to control bipedal robots to achieve flexible omni-directional motion and robust locomotion under complex disturbances, called the heuristics-based reinforcement learning (HBRL) framework. HBRL shows great training efficiency in simulation. Heuristic reference trajectories play a crucial role in HBRL, which guide the training process. Exploration rewards, leg-foot reset condition, and command curriculum are three significant components to optimize the training process. An estimator network is utilized to supply linear velocities and foot contact information. We train controllers on flat ground in simulation. To demonstrate robustness and versatility, the trained controllers were tested on BRAVER, a point-foot bipedal robot with three joints on each leg. The controllers enabled BRAVER to perform omni-directional locomotion with the maximum forward speed reaching 2[Formula: see text]m/s. The robot could also maintain balance under external pushing and over uneven terrains.
科研通智能强力驱动
Strongly Powered by AbleSci AI