强化学习
马尔可夫决策过程
运动规划
机器人
计算机科学
任务(项目管理)
过程(计算)
自动化
路径(计算)
人工智能
理论(学习稳定性)
机器人学
部分可观测马尔可夫决策过程
功能(生物学)
移动机器人
机器人控制
控制(管理)
机器人学习
工程类
模拟
自主机器人
自动计划和调度
马尔可夫过程
控制工程
决策过程
作者
Shougang Sui,Zhiqi Liu,Jiabin Huang,Shizhan Luo,Xiuzhe Liu,Kun Li
标识
DOI:10.1177/01423312261429128
摘要
With the rapid development of robotic and artificial intelligence technologies, the autonomous decision-making capability of endoscopic surgical robot has been significantly enhanced, accompanied by growing demands for automation in non-decision-making tasks. This study focuses on the path planning for complex tasks in surgical robot by proposing a hierarchical reinforcement learning (HRL) framework based on the option framework. Within this framework, the Semi-Markov Decision Process (SMDP) is extended into an augmented Markov Decision Process (MDP) to optimize termination conditions and facilitate long-horizon task training. To address the sparse reward problem, a hierarchical reward function is designed with intrinsic temporal rewards specifically implemented for high-level policies. In addition, a Type-Shared Option Policy (TSOP) is proposed to enhance training efficiency. Experimental results demonstrate that the proposed HRL framework effectively improves both the success rate and stability of path planning for surgical robot in the da Vinci Research Kit (dVRK) simulation environment.
科研通智能强力驱动
Strongly Powered by AbleSci AI