计算机科学
运动规划
马尔可夫决策过程
移动机器人
强化学习
路径(计算)
运动学
人工智能
过程(计算)
机器人
自主代理人
分布式计算
机器学习
实时计算
马尔可夫过程
计算机网络
统计
经典力学
操作系统
物理
数学
作者
Lixiang Zhang,Ze Cai,Yan Yan,Chen Yang,Yaoguang Hu
标识
DOI:10.1016/j.engappai.2023.107631
摘要
The study addresses path planning problems for autonomous mobile robots (AMRs), considering their kinematics, where performance and responsiveness are often incompatible. This study proposes a multi-agent policy learning-based method to tackle this challenge in dynamic environments. The proposed method features a centralized learning and decentralized execution-based path planning framework designed to meet performance and responsiveness requirements. The problem is modeled as a partial observation Markov Decision Process for policy learning while considering the kinematics using conventional neural networks. Then, an improved proximal policy optimization algorithm is developed with highlight experience replay that corrects failed experiences to speed up the learning processes. The experimental results show that the proposed method outperforms the baselines in both static and dynamic environments. The proposed method shortens the movement distance and time in static environments by about 29.1% and 5.7%, as well as in dynamic environments by about 21.1% and 20.4%, respectively. The runtime is maintained in milliseconds across various environments, taking only 0.07 s. Overall, the proposed method is valid and efficient in ensuring the performance and responsiveness of AMRs when dealing with complex and dynamic path planning problems.
科研通智能强力驱动
Strongly Powered by AbleSci AI