强化学习
困境
社会困境
动作(物理)
计算机科学
人工智能
博弈论
社会学习
进化计算
钢筋
数理经济学
心理学
数学
社会心理学
知识管理
量子力学
物理
几何学
作者
Litong Fan,Dengxiu Yu,Zhen Wang
标识
DOI:10.1109/tcss.2024.3409833
摘要
This article presents a framework for exploring optimal evolutionary strategies in continuous-action social dilemma games with a hierarchical structure comprising a leader and multifollowers. Previous studies in game theory have frequently overlooked the hierarchical structure among individuals, assuming that decisions are made simultaneously. Here, we propose a hierarchical structure for continuous action games that involves a leader and followers to enhance cooperation. The optimal evolutionary strategy for the leader is to guide the followers’ actions to maximize overall benefits by exerting minimal control, while the followers aim to maximize their payoff by making minimal changes to their strategies. We establish the coupled Hamilton–Jacobi–Bellman (HJB) equations to find the optimal evolutionary strategy. To address the complexity of asymmetric roles arising from the leader-follower structure, we introduce an integral reinforcement learning (RL) algorithm known as two-level heuristic dynamic programming (HDP)-based value iteration (VI). The implementation of the algorithm utilizes neural networks (NNs) to approximate the value functions. Moreover, the convergence of the proposed algorithm is demonstrated. Additionally, three social dilemma models are presented to validate the efficacy of the proposed algorithm.
科研通智能强力驱动
Strongly Powered by AbleSci AI