强化学习
计算机科学
理论(学习稳定性)
控制(管理)
风速
人工智能
功能(生物学)
控制理论(社会学)
模拟
机器学习
地理
进化生物学
生物
气象学
作者
Jun Xue,Ziniu Liu,Guanjun Liu,Ziyuan Zhou,Kaiwen Zhang,Ying Tang,Jiacun Wang
标识
DOI:10.1109/tiv.2023.3324687
摘要
Unmanned Aerial Vehicles (UAVs) have extensive applications such as logistics transportation and aerial photography. However, UAVs are sensitive to winds. Traditional control methods, such as proportional- integral-derivative controllers, generally fail to work well when the strength and direction of winds are changing frequently. In this work deep reinforcement learning algorithms are combined with a domain randomization method to learn robust wind-resistant hovering policies. A novel reward function is designed to guide learning. This reward function uses a constant reward to maintain a continuous flight of a UAV as well as a weight of the horizontal distance error to ensure the stability of the UAV at altitude. A five-dimensional representation of actions instead of the traditional four dimensions is designed to strengthen the coordination of wings of a UAV. We theoretically explain the rationality of our reward function based on the theories of Q-learning and reward shaping. Experiments in the simulation and real-world application both illustrate the effectiveness of our method. To the best of our knowledge, it is the first paper to use reinforcement learning and domain randomization to explore the problem of robust wind-resistant hovering control of quadrotor UAVs, providing a new way for the study of wind-resistant hovering and flying of UAVs.
科研通智能强力驱动
Strongly Powered by AbleSci AI