乘法函数
直觉
趋同(经济学)
计算机科学
数学
虚构的游戏
收敛速度
正规化(语言学)
数学优化
算法
延缓
数理经济学
向前看
应用数学
战略
存在量化
学习效果
作者
Yang Cai,Gabriele Farina,Julien Grand-Clément,Christian Kroer,Chung-Wei Lee,Haipeng Luo,Weiqiang Zheng
出处
期刊:Operations Research
[Institute for Operations Research and the Management Sciences]
日期:2026-08-14
标识
DOI:10.1287/opre.2025.2166
摘要
Convergence of Optimistic Learning in Games and the Role of Forgetfulness Online learning algorithms solve games by repeatedly updating players’ strategies. A natural hope is that the latest strategy improves at a predictable rate. This paper shows that this intuition can fail for optimistic follow-the-regularized-leader methods, including the widely used optimistic multiplicative weights algorithm. In two-player zero-sum games, the authors separate three notions of nonergodic performance: last-iterate, random-iterate, and best-iterate convergence. They prove that no instance-independent last-iterate rate exists for the broad algorithmic family and establish strong lower bounds for random iterates. Yet, a useful positive result remains; in 2 × 2 games, optimistic multiplicative weights achieve a uniform best-iterate rate. The analysis traces the slowdown to a lack of “forgetfulness”; accumulated past losses can keep the dynamics moving away from equilibrium long after they reach its neighborhood. The results suggest that practitioners should distinguish carefully between the latest strategy, a randomly selected strategy, and the best observed strategy when evaluating learning dynamics.
科研通智能强力驱动
Strongly Powered by AbleSci AI