可观测性
强化学习
计算机科学
人工智能
深度学习
机器人学
噪音(视频)
特征工程
特征(语言学)
循环神经网络
人工神经网络
机器人
图像(数学)
数学
哲学
语言学
应用数学
作者
Godwyll Aikins,Sagar Jagtap,Kim‐Doang Nguyen
出处
期刊:Drones
[Multidisciplinary Digital Publishing Institute]
日期:2024-05-30
卷期号:8 (6): 232-232
被引量:15
标识
DOI:10.3390/drones8060232
摘要
Landing a multi-rotor uncrewed aerial vehicle (UAV) on a moving target in the presence of partial observability, due to factors such as sensor failure or noise, represents an outstanding challenge that requires integrative techniques in robotics and machine learning. In this paper, we propose embedding a long short-term memory (LSTM) network into a variation of proximal policy optimization (PPO) architecture, termed robust policy optimization (RPO), to address this issue. The proposed algorithm is a deep reinforcement learning approach that utilizes recurrent neural networks (RNNs) as a memory component. Leveraging the end-to-end learning capability of deep reinforcement learning, the RPO-LSTM algorithm learns the optimal control policy without the need for feature engineering. Through a series of simulation-based studies, we demonstrate the superior effectiveness and practicality of our approach compared to the state-of-the-art proximal policy optimization (PPO) and the classical control method Lee-EKF, particularly in scenarios with partial observability. The empirical results reveal that RPO-LSTM significantly outperforms competing reinforcement learning algorithms, achieving up to 74% more successful landings than Lee-EKF and 50% more than PPO in flicker scenarios, maintaining robust performance in noisy environments and in the most challenging conditions that combine flicker and noise. These findings underscore the potential of RPO-LSTM in solving the problem of UAV landing on moving targets amid various degrees of sensor impairment and environmental interference.
科研通智能强力驱动
Strongly Powered by AbleSci AI