Deep Reinforcement Learning for Nash Equilibrium of Differential Games

纳什均衡 强化学习 ε平衡 最佳反应 计算机科学 极小极大 数学优化 相关平衡 梯度下降 均衡选择 博弈论 数理经济学 人工智能 数学 重复博弈 人工神经网络
作者
Zhenyu Li,Ya-Zhong Luo
出处
期刊:IEEE transactions on neural networks and learning systems [Institute of Electrical and Electronics Engineers]
卷期号:36 (2): 2747-2761 被引量:22
标识
DOI:10.1109/tnnls.2024.3351631
摘要

Nash equilibrium is a significant solution concept representing the optimal strategy in an uncooperative multiagent system. This study presents two deep reinforcement learning (DRL) algorithms for solving the Nash equilibrium of differential games. Both algorithms are built upon the distributed distributional deep deterministic policy gradient (D4PG) algorithm, which is a one-sided learning method. We modified it to a two-sided adversarial learning method. The first is D4PG for games (D4P2G), which directly applies an adversarial play framework based on the D4PG. A simultaneous policy gradient descent (SPGD) method is employed to optimize the policies of the players with conflicting objectives. The second is the distributional deep deterministic symplectic policy gradient (D4SPG) algorithm, which is our main contribution. More specifically, it newly designs a minimax learning framework that combines the critics of the two players and proposes a symplectic policy gradient adjustment method to find a better policy gradient. Simulations show that both algorithms converge to the Nash equilibrium in most cases, but D4SPG can learn the Nash equilibrium more accurately and efficiently, especially in Hamiltonian games. Moreover, it can handle games with complex dynamics, which is challenging for traditional methods.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
刚刚
刚刚
共享精神的应助被BCZ采纳,获得10
刚刚
刚刚
刚刚
刚刚
刚刚
LarryC完成签到,获得积分10
1秒前
科研通AI6.4的应助被董凡侨采纳,获得10
1秒前
幸福胡萝卜完成签到,获得积分10
1秒前
华风发布了新的文献求助10
2秒前
Summer发布了新的文献求助10
2秒前
离个大谱发布了新的文献求助10
2秒前
胡咩咩完成签到,获得积分10
2秒前
慕苡发布了新的文献求助10
3秒前
lanrete完成签到,获得积分10
3秒前
3秒前
123asd发布了新的文献求助10
3秒前
艾斯福完成签到,获得积分10
3秒前
木尧完成签到,获得积分20
4秒前
mimi完成签到,获得积分10
4秒前
居居不酷完成签到,获得积分10
5秒前
俊逸的平卉完成签到 ,获得积分10
5秒前
科研顺利完成签到,获得积分10
5秒前
科研小狗发布了新的文献求助10
6秒前
511发布了新的文献求助10
6秒前
小周完成签到 ,获得积分10
6秒前
lumos发布了新的文献求助10
7秒前
Maqian发布了新的文献求助10
7秒前
trz817394发布了新的文献求助10
8秒前
Nole的应助被小邓采纳,获得10
8秒前
皮夏寒发布了新的文献求助10
8秒前
天邪完成签到,获得积分10
8秒前
SK完成签到,获得积分10
8秒前
123给123的求助进行了留言
8秒前
月亮门儿发布了新的文献求助10
9秒前
段段完成签到,获得积分10
9秒前
科研通AI6.4的应助被HHHHHN采纳,获得10
10秒前
阿落发布了新的文献求助10
10秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
Organizational Behavior 510
Management and the Arts 510
Geschichtliche Grundbegriffe (GGB), Band 5: Pro–Soz 300
Die Religion in Geschichte und Gegenwart (RGG), 4. Auflage, Band 7: R–S 300
Die Religion in Geschichte und Gegenwart (RGG), 4. Auflage, Band 1: A–B 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 内科学 物理 有机化学 化学工程 生物化学 复合材料 光电子学 细胞生物学 心理学 量子力学 催化作用 物理化学 电极
热门帖子
关注 科研通微信公众号,转发送积分 7793307
求助须知:如何正确求助?哪些是违规求助? 9329972
关于积分的说明 20434132
捐赠科研通 7383265
什么是DOI,文献DOI怎么找? 3323956
关于科研通互助平台的介绍 2471829
邀请新用户注册赠送积分活动 2341026