强化学习
马尔可夫决策过程
计算机科学
启发式
地铁列车时刻表
调度(生产过程)
人工神经网络
数学优化
最优化问题
运筹学
人工智能
马尔可夫过程
工程类
统计
操作系统
数学
算法
作者
Cheng-shuo Ying,Andy H.F. Chow,Yihui Wang,Kwai‐Sang Chin
标识
DOI:10.1109/tits.2021.3063399
摘要
This paper presents an integrated metro service scheduling and train unit deployment with a proximal policy optimization approach based on the deep reinforcement learning framework. The optimization problem is formulated as a Markov decision process (MDP) subject to a set of operational constraints. To address the computational complexity, the value function and control policy are parameterized by artificial neural networks (ANNs) with which the operational constraints are incorporated through a devised mask scheme. A proximal policy optimization (PPO) approach is developed for training the ANNs via successive transition simulations. The optimization framework is implemented and tested on a real-world scenario configured with the Victoria Line of London Underground, UK. The results show that the performance of proposed methodology outperforms a set of selected evolutionary heuristics in terms of both solution quality and computational efficiency. Results illustrate the advantages of having flexible train composition in saving operational costs and reducing service irregularities. This study contributes to real time metro operations with limited resources and state-of-art optimization techniques.
科研通智能强力驱动
Strongly Powered by AbleSci AI