排
强化学习
计算机科学
初始化
功能(生物学)
趋同(经济学)
人工智能
编码(集合论)
弹道
机器学习
任务(项目管理)
理论(学习稳定性)
分布式计算
模式
任务分析
协调博弈
进化计算
进化算法
作者
Dixiao Wei,Peng Yi,Jinlong Lei,Yiguang Hong,Hairong Dong,Yuchuan Du
标识
DOI:10.1109/tits.2026.3664858
摘要
Reinforcement Learning (RL) has demonstrated excellent decision-making potential in platoon coordination problems. However, due to the variability of coordination goals, the complexity of the decision problem, and the time-consumption of trial-and-error in manual design, finding a well performance reward function to guide RL training to solve complex platoon coordination problems remains challenging. In this paper, we formally define the Platoon Coordination Reward Design Problem (PCRDP), extending the RL-based cooperative platoon coordination problem to incorporate automated reward function generation. To address PCRDP, we propose a Large Language Model (LLM)-based Platoon coordination Reward Design (PCRD) framework, which systematically automates reward function discovery through LLM-driven initialization and iterative optimization. In this method, LLM first initializes reward functions based on environment code and task requirements with an Analysis and Initial Reward (AIR) module, and then iteratively optimizes them based on training feedback with an evolutionary module. The AIR module guides LLM to deepen their understanding of code and tasks through a chain of thought, effectively mitigating hallucination risks in code generation. The evolutionary module fine-tunes and reconstructs the reward function, achieving a balance between exploration diversity and convergence stability for training. To validate our approach, we establish six challenging coordination scenarios with varying complexity levels within the Yangtze River Delta transportation network simulation. Comparative experimental results demonstrate that RL agents utilizing PCRD-generated reward functions consistently outperform human-engineered reward functions, achieving an average of 10% higher performance metrics in all scenarios.
科研通智能强力驱动
Strongly Powered by AbleSci AI