清晨好,您是今天最早来到科研通的研友!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您科研之路漫漫前行!

Path-Following Control of Unmanned Underwater Vehicle Based on an Improved TD3 Deep Reinforcement Learning

强化学习 遥控水下航行器 水下 钢筋 控制(管理) 计算机科学 路径(计算) 控制理论(社会学) 工程类 移动机器人 人工智能 机器人 计算机网络 结构工程 海洋学 地质学
作者
Yexin Fan,Hongyang Dong,Xiaowei Zhao,Petr Denissenko
出处
期刊:IEEE Transactions on Control Systems and Technology [Institute of Electrical and Electronics Engineers]
卷期号:32 (5): 1904-1919 被引量:50
标识
DOI:10.1109/tcst.2024.3377876
摘要

This work proposes an innovative path-following control method, anchored in deep reinforcement learning (DRL), for unmanned underwater vehicles (UUVs). This approach is driven by several new designs, all of which aim to enhance learning efficiency and effectiveness and achieve high-performance UUV control. Specifically, a novel experience replay strategy is designed and integrated within the twin-delayed deep deterministic policy gradient algorithm (TD3). It distinguishes the significance of stored transitions by making a trade-off between rewards and temporal-difference (TD) errors, thus enabling the UUV agent to explore optimal control policies more efficiently. Another major challenge within this control problem arises from action oscillations associated with DRL policies. This issue leads to excessive system wear on actuators and makes real-time application difficult. To mitigate this challenge, a newly improved regularization method is proposed, which provides a moderate level of smoothness to the control policy. Furthermore, a dynamic reward function featuring adaptive constraints is designed to avoid unproductive exploration and expedite learning convergence speed further. Simulation results show that our method garners higher rewards in fewer training episodes compared with mainstream DRL-based control approaches (e.g., deep deterministic policy gradient (DDPG) and vanilla TD3) in UUV applications. Moreover, it can adapt to varying path configurations amid uncertainties and disturbances, all while ensuring high tracking accuracy. Simulation and experimental studies are conducted to verify the effectiveness.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
追寻书本完成签到,获得积分10
6秒前
30秒前
漂亮忆曼完成签到,获得积分10
38秒前
明亮访梦完成签到,获得积分10
52秒前
56秒前
王志新完成签到 ,获得积分10
57秒前
美好的初翠完成签到,获得积分10
59秒前
wxh完成签到 ,获得积分10
1分钟前
时尚靖琪完成签到,获得积分10
1分钟前
欣喜的涵柏完成签到 ,获得积分10
1分钟前
Lan完成签到 ,获得积分10
1分钟前
酷波er的应助被fcyyc采纳,获得10
1分钟前
大模型的应助被科研通管家采纳,获得10
1分钟前
1分钟前
fcyyc完成签到,获得积分10
1分钟前
1分钟前
zhangsan发布了新的文献求助10
1分钟前
靓丽雪萍完成签到,获得积分10
1分钟前
如意夜云完成签到,获得积分10
1分钟前
怕黑的驳完成签到,获得积分10
2分钟前
fcyyc发布了新的文献求助10
2分钟前
zhangsan完成签到,获得积分10
2分钟前
傲娇的期待完成签到,获得积分10
2分钟前
专注凌柏完成签到,获得积分10
2分钟前
华仔的应助被eas采纳,获得10
2分钟前
RYYYYYYY233完成签到 ,获得积分10
2分钟前
2分钟前
2分钟前
感性的雅霜完成签到,获得积分10
2分钟前
yzyyzy发布了新的文献求助10
3分钟前
风中的晓兰完成签到,获得积分10
3分钟前
陈M雯完成签到 ,获得积分10
3分钟前
tao完成签到 ,获得积分10
3分钟前
3分钟前
aaaaa发布了新的文献求助10
3分钟前
虚无完成签到,获得积分10
3分钟前
飞云完成签到 ,获得积分10
3分钟前
傻傻的冰烟完成签到,获得积分10
3分钟前
YifanWang的应助被科研通管家采纳,获得10
3分钟前
悦耳的白云完成签到,获得积分10
3分钟前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
Computational Chemical Reaction Engineering: Modeling, Simulation, and Design with MATLAB 600
Organizational Behavior 510
Management and the Arts 510
Production Logging: Theoretical and Interpretive Elements 400
CLSI C56QG Examples of Hemolyzed, Icteric, and Lipemic/Turbid Samples Quick Guide 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 内科学 物理 有机化学 化学工程 生物化学 复合材料 光电子学 细胞生物学 心理学 量子力学 催化作用 物理化学 电极
热门帖子
关注 科研通微信公众号,转发送积分 7817008
求助须知:如何正确求助?哪些是违规求助? 9345751
关于积分的说明 20531090
捐赠科研通 7409382
什么是DOI,文献DOI怎么找? 3331611
关于科研通互助平台的介绍 2477819
邀请新用户注册赠送积分活动 2351211