反事实思维
计算机科学
任务(项目管理)
强化学习
质量(理念)
控制(管理)
人工智能
动作(物理)
机器学习
分配问题
增强学习
机器人
机制(生物学)
钥匙(锁)
任务分析
多智能体系统
作者
Sitong Qian,Meibao Yao,Xueming Xiao
摘要
In multi-agent tasks, the credit assignment problem has garnered widespread attention, as an effective credit assignment mechanism can not only promote cooperation among agents but also significantly enhance the efficiency and quality of task completion. Reasonable credit assignment ensures that each agent receives its due rewards, thus motivating them to participate and coordinate more actively, which is crucial for successfully completing complex tasks. To address this issue, we propose an improved algorithm based on the MAPPO algorithm. This algorithm employs a centralized training and decentralized execution framework, incorporates a counterfactual baseline, and uses an enhanced action value network as a centralized critic. By fixing the behaviors of other agents and marginalizing the actions of a single agent, we can more accurately assess each agent's contribution to the overall team performance. This approach not only ensures the fairness of credit assignment but also allows for a granular assessment of each agent's specific performance, achieving a more refined reward mechanism. To validate the effectiveness of our improved algorithm, we conducted experiments in a collaborative robotic navigation task. The results indicate that our method outperforms existing approaches in overall performance.
科研通智能强力驱动
Strongly Powered by AbleSci AI