强化学习
计算机科学
钢筋
控制(管理)
信号(编程语言)
人工智能
工程类
结构工程
程序设计语言
作者
Hao Huang,Hang Qi,Zhiqun Hu,Qilie Liu,Qian Liu,Xiaohua Xu,Hang Li
出处
期刊:IEEE Transactions on Vehicular Technology
[Institute of Electrical and Electronics Engineers]
日期:2025-07-31
卷期号:75 (1): 233-249
标识
DOI:10.1109/tvt.2025.3594343
摘要
With the advancement of artificial intelligence, Deep Reinforcement Learning (DRL) has been widely applied to traffic signal control (TSC). However, DRL-based TSC methods often require extensive training samples and struggle to generalize across heterogeneous traffic environments, such as different intersection types and traffic flows. Retraining the DRL agent for new environments inevitably increases both costs and time investments, limiting the flexibility of these intelligent TSC algorithms in practical applications. In this paper, we propose an inductive meta-DRL-based traffic signal control method, named IM-TSC, for heterogeneous environments to address these issues. First, we design a feature extraction module composed of a self-feature induction network and a neighborhood feature extraction network to obtain accurate and comprehensive traffic representations, laying the foundation for generalization and coordination. Furthermore, we decouple the scenario inference and agent training processes and design a scenario inference network to obtain contextual encoding for the current intersection scenario through contrastive learning, thereby promoting the transfer of meta-knowledge across similar scenarios. Finally, we design a meta-training framework for the DRL agent with a continuous action space, enabling efficient training of the meta-learner by alternating between local-level learning and global-level learning. We evaluate the proposed IM-TSC method across multiple heterogeneous traffic environments. Experimental results demonstrate that the meta-learner trained with the IM-TSC method can adapt to heterogeneous environments more quickly and effectively. Compared to the suboptimal baseline, it achieves average reductions of 9.20% in the number of stops, 11.55% in waiting time, and 8.29% in fuel consumption.
科研通智能强力驱动
Strongly Powered by AbleSci AI