进化博弈论
强化学习
趋同(经济学)
计算机科学
透视图(图形)
博弈论
动力系统理论
进化动力学
进化算法
数学优化
人工智能
纳什均衡
潜在博弈
进化规划
理论计算机科学
动力学(音乐)
复制因子方程
进化计算
类型(生物学)
系统动力学
动力系统(定义)
人工生命
数学
复杂系统
机器学习
作者
J. Bauer,Sheldon West,Eduardo Alonso,Mark Broom
标识
DOI:10.1098/rspa.2025.0449
摘要
Abstract We pursue a mathematically rigorous approach connecting stochastic multi-agent reinforcement learning (MARL) processes to deterministic dynamical systems of replicator–mutator dynamics (RMD) type from evolutionary game theory (EGT). This dynamical systems perspective makes the rich literature on evolutionary game dynamics directly available for establishing theoretical guarantees for algorithm convergence in complex multi-agent environments, addressing a fundamental challenge in the field. We demonstrate this approach by presenting and analysing mutation-bias learning with direct policy updates (MBL-DPU), a MARL algorithm that provably approximates RMD, and show convergence in stable games. Through experiments across games of increasing complexity and dimensionality, we demonstrate our dynamical systems analysis and the convergence of MBL-DPU while win-or-learn-fast policy-hill-climbing (WoLF-PHC) and frequency-adjusted Q-learning (FAQ), algorithms with fewer theoretical guarantees, deteriorate unexpectedly in higher dimensions. Beyond specific algorithms, the approach demonstrates a principled route to transferring results from EGT to multi-agent learning, allowing a systematic comparison of evolutionary game dynamics with MARL algorithms and enabling the derivation of further MARL algorithms from evolutionary dynamics. This approach further questions the assumption that algorithm complexity generally improves performance and underscores the necessity of mathematical rigour in analysing MARL algorithms. For a systematic comparison, we also introduce and experimentally analyse mutationbias learning with logistic choice (MBL-LC), a variant closer to Q-learning but lacking the theoretical guarantees of MBL-DPU.
科研通智能强力驱动
Strongly Powered by AbleSci AI