后悔
计算机科学
区间(图论)
理论(学习稳定性)
上下界
数学优化
弹道
纳什均衡
在线算法
在线学习
古诺竞争
付款
重复博弈
开放集
游戏中的回合、回合和计时系统
序贯博弈
非合作博弈
数理经济学
数学
博弈论
算法
虚构的游戏
指数稳定性
筛选游戏
机构设计
极大极小定理
作者
Wenting Liu,Jinlong Lei,Peng Yi,Lacra Pavel
标识
DOI:10.1109/tac.2025.3639124
摘要
This paper considers online learning for open non-cooperative games where players can join and leave the system freely, while the current number of players in the system and the opponents' identities are not available. Unlike existing works on closed non-cooperative games that assume a fixed number of players, the open scenario setting is characterized by time-varying payment functions as well as a time-varying number of players. We present an online learning mechanism based on the best-response algorithm that enables players to adaptively adjust their strategies in the open game with anonymous opponents, thereby minimizing their own payments. We first provide an upper bound on the adaptive dynamic regret of the algorithm, which measures each player's regret over the time interval from joining to leaving the system. Then, we prove that the open game system is open stable with a stability radius$R$, where$R$depends on the time-variation of the equilibrium trajectory as well as on the ratios of newly joined and departing players to the number of active players. Finally, we demonstrate the algorithm performance through numerical simulations on an open Cournot game.
科研通智能强力驱动
Strongly Powered by AbleSci AI