Deep Reinforcement Learning With NMPC Assistance Nash Switching for Urban Autonomous Driving

强化学习 钢筋 计算机科学 人工智能 心理学 社会心理学
作者
Sina Alighanbari,Nasser L. Azad
出处
期刊:IEEE transactions on intelligent vehicles [Institute of Electrical and Electronics Engineers]
卷期号:8 (3): 2604-2615 被引量:27
标识
DOI:10.1109/tiv.2022.3167616
摘要

Deep Deterministic Policy Gradient (DDPG) is a promising reinforcement learning technique with the potential to resolve complicated tasks and handle high-dimensional state/action spaces. However, it suffers from sample inefficiency, requiring a high number of training samples. To speed up the training, we propose$\epsilon$-annealing and Q-learning switching methods to aid the training of DDPG with Nonlinear Model Predictive Control (NMPC) controller to solve priority calculation and merging of autonomous vehicles at roundabouts. We further expand the Q-learning switch with double replay memory and Nash Q-value updates. The performance of these switching methods are compared to DDPG and demonstrate that Nash switch outperforms other methods. To reduce conservativeness, we test training using variable traffic density. We test three selection methods inside Q-learning and show constant threshold switch has at least ten times higher mean reward for 50 episodes training. We also compare Q-learning with NMPC and PID assistance and show that NMPC has 114% higher mean reward. We compare Q-learning switch and novel Nash switch method under noise-free and noisy input conditions to prove an increase of 35% mean reward and decrease of 4% std for Nash updates. We analyze efficacy of Q-learning and Nash switch approaches w.r.t NMPC and demonstrate comparable performance between Nash switch and NMPC. We juxtapose driving results of switch Q-learning and Nash switch with DDPG algorithm to prove Nash switch strategy has higher overall performance. Finally, we compare Nash switch’s performance with DDPG for highway merging scenario which shows 159% higher mean reward.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
怕黑岱周发布了新的文献求助10
1秒前
怕黑岱周发布了新的文献求助30
1秒前
壮观溪流发布了新的文献求助10
1秒前
怕黑岱周发布了新的文献求助10
1秒前
ding应助忧郁翠彤采纳,获得10
1秒前
完美世界应助myg8627采纳,获得10
2秒前
2秒前
FashionBoy应助泡芙采纳,获得10
2秒前
3秒前
3秒前
怕黑岱周发布了新的文献求助30
3秒前
夕夕发布了新的文献求助20
4秒前
初景发布了新的文献求助10
4秒前
怕黑岱周发布了新的文献求助10
4秒前
一颗西米子完成签到,获得积分10
4秒前
4秒前
万能图书馆应助haoqingyun采纳,获得10
4秒前
Anwz完成签到,获得积分10
4秒前
4秒前
怕黑岱周发布了新的文献求助10
5秒前
怕黑岱周发布了新的文献求助10
5秒前
怕黑岱周发布了新的文献求助10
5秒前
怕黑岱周发布了新的文献求助10
5秒前
5秒前
5秒前
5秒前
6秒前
6秒前
不一样的烟火完成签到,获得积分10
6秒前
Lucas应助旺旺仙贝采纳,获得10
7秒前
7秒前
榕树发布了新的文献求助10
7秒前
7秒前
8秒前
可ke完成签到,获得积分10
8秒前
小王发布了新的文献求助10
8秒前
8秒前
怕黑岱周发布了新的文献求助10
9秒前
怕黑岱周发布了新的文献求助10
9秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
化工安全与环保 1000
Autoparametric Resonance in Mechanical Systems 1000
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 800
Effects of Two Weeks of Red Light Therapy on Choroidal Thickness and Axial Length in Young Adults 700
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 600
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7658117
求助须知:如何正确求助?哪些是违规求助? 9228579
关于积分的说明 19836963
捐赠科研通 7224837
什么是DOI,文献DOI怎么找? 3280790
关于科研通互助平台的介绍 2440790
邀请新用户注册赠送积分活动 2280579