避障
情态动词
强化学习
计算机科学
障碍物
钢筋
移动机器人
人工智能
声学
机器人
结构工程
工程类
材料科学
物理
地理
考古
高分子化学
作者
Zhaoqing Lu,Li He,Hongwei Wang,Liang Yuan,Wendong Xiao,Zhening Liu,Yao‐Hua Chen
标识
DOI:10.1088/1361-6501/adb1ff
摘要
Abstract The multimodal perception is crucial for mobile robots to achieve safe autonomous steering in complex and changing environments without colliding with obstacles. In spite of the fact that rich environmental information can be provided by the multimodal data, how to integrate the complementary features from different sources effectively remains an open issue. In this paper, a simple yet effective multimodal deep reinforcement learning (DRL) method, concise and effective multimodal DRL, by fusing two input modalities such as depth images and pseudo-LiDAR data generated by RGB-D cameras, is proposed to improve the performance of single sensor in environmental perception tasks. And the convolutional LSTM network layers are also adopted to extract spatiotemporal features and capture temporal relationships in consecutive time steps, which allow the robot to predict dynamic changes of the environment. In addition, A cross-modal fusion module is designed to optimize the multimodal data fusion scheme, with dynamic weighting and the merging of features from different modalities, allowing the agent to focus more reasonably on the current state and improving the accuracy and efficiency of the obstacle avoidance strategy. The experimental results show that the proposed approach achieved significant performance in cumulative rewards, convergence speed, and success rate. Additionally, tests in both simulated and real-world environments further verify that collisions are effectively avoided using only a low-cost RGB-D camera, demonstrating the method’s strong generalization capabilities.
科研通智能强力驱动
Strongly Powered by AbleSci AI