
浏览全部资源
扫码关注微信
1.中国铁道科学研究院,北京 100081
2.中国铁路列车运行图技术中心,北京 100081
3.中国铁道科学研究院集团有限公司 运输及经济研究所,北京 100081
4.中国铁道学会,北京 100081
周黎(1963—),男,四川广元人,正高级工程师,博士,从事铁路运输组织研究;E-mail:Zl.mor@163.com
收稿:2025-09-18,
网络首发:2026-07-24,
纸质出版:2026-07-28
移动端阅览
王岩,周黎,范家铭等.基于深度强化学习的大型高铁站到发线运用方案快速编制方法[J].铁道科学与工程学报,2026,23(07):3099-3109.
WANG Yan,ZHOU Li,FAN Jiaming,et al.A fast approach for station track utilization at large high-speed railway station based on deep reinforcement learning[J].Journal of Railway Science and Engineering,2026,23(07):3099-3109.
王岩,周黎,范家铭等.基于深度强化学习的大型高铁站到发线运用方案快速编制方法[J].铁道科学与工程学报,2026,23(07):3099-3109. DOI: 10.19713/j.cnki.43-1423/u.T20251484.
WANG Yan,ZHOU Li,FAN Jiaming,et al.A fast approach for station track utilization at large high-speed railway station based on deep reinforcement learning[J].Journal of Railway Science and Engineering,2026,23(07):3099-3109. DOI: 10.19713/j.cnki.43-1423/u.T20251484.
大型高铁站是高铁网络中的关键枢纽节点,其到发线运用方案的合理安排是车站运输组织工作的核心。为实现大型高铁站到发线运用方案的快速编制,将人工智能领域的深度强化学习方法应用于到发线运用问题中。首先,基于列车过站径路的概念来实现到发线分配与进路排列的协同优化,将到发线运用问题描述为最大独立集问题,以尽可能安排所有列车为目标,形成整数规划模型。然后,构建基于冲突图的强化学习环境模型,将到发线运用方案数学模型转化为马尔可夫决策过程,定义状态表征、动作空间、状态转移和奖励函数,为智能体提供训练数据的交互采集场所。最后,设计基于GCN(graph convolutional network)的到发线运用智能体神经网络架构,基于PPO(proximal policy optimization)算法实现智能体高效预训练,在决策阶段,设计基于2Imp的决策提升算法进一步优化方案编制质量。以北京南高速场和广州南站为案例场景,对比3种不同方法的求解时间与最终目标函数取值。结果表明,经过4.2 h预训练后智能体完成收敛,决策阶段智能算法可以快速求得最优解,相比于运筹学经典算法,求解质量相同的情况下智能算法求解速度可提升5~16倍,1.2 s即可完成广州南站1 133列列车的到发线运用方案编制。此外,该方法具有良好的扩展性和泛化能力,智能体经过一次训练后即可泛化至不同车站、不同规模的场景中进行决策与求解。研究结果可为进一步提升列车运行图整体编制效率提供技术支撑,为铁路运输组织的智能化发展提供方法参考。
Large high-speed railway stations serve as critical hub nodes in high-speed rail networks
where the rational arrangement of station track utilization plans constitutes the core of station transportation organization. To enable rapid generation of track utilization schemes for large high-speed stations
this study proposed applying deep reinforcement learning from artificial intelligence to address track utilization problems. Firstly
based on the concept of train passing routes
this paper achieved collaborative optimization of track allocation and route arrangement. The track utilization problem was formulated as a maximum independent set problem with the objective of accommodating all trains
leading to an integer programming model. Subsequently
a conflict-graph-based reinforcement learning environment model was constructed to transform the mathematical model of track utilization into a Markov Decision Process. This paper defined the state representation
action space
state transition
and reward function to establish an interactive data collection framework for agent training. Finally
a graph convolutional network (GCN)-based intelligent agent architecture was designed
combined with proximal policy optimization (PPO) algorithm for efficient pre-training. During the decision phase
a 2Imp-based decision enhancement algorithm was proposed to improve solution quality. The case studies of Beijing South High-Speed Field and Guangzhou South Station were conducted to compare the solving time and final objective function values of three different methods. The results demonstrate that after 4.2 hours of pre-training
the agent achieves convergence and rapidly obtains optimal solutions during the decision phase. Compared to classical operations research algorithms
the proposed intelligent algorithm can achieve equivalent solution quality with 5 to 16 times faster
completing the track utilization planning for 1 133 trains at Guangzhou South Station within 1.2 seconds. Moreover
the proposed method can exhibit strong scalability and transferability
enabling the agent to be directly deployed to stations of different scales and configurations for decision-making and problem-solving after a single training phase. These findings can provide technical support for enhancing the overall efficiency of train timetable preparation. The results can offer methodological references for the intelligent development of railway transportation organization.
ZWANEVELD P J , KROON L G , ROMEIJN H E , et al . Routing trains through railway stations: model formulation and algorithms [J ] . Transportation Science , 1996 , 30 ( 3 ): 181 - 194 .
BILLIONNET A . Using integer programming to solve the train-platforming problem [J ] . Transportation Science , 2003 , 37 ( 2 ): 213 - 222 .
高全 , 张英贵 , 陈治亚 , 等 . 考虑冗余时间的区域铁路车站股道运用计划协同编制优化方法 [J ] . 铁道科学与工程学报 , 2023 , 20 ( 2 ): 506 - 515 .
GAO Quan , ZHANG Yinggui , CHEN Zhiya , et al . Track utilization plan coordinative optimal method for regional railway stations based on redundant time [J ] . Journal of Railway Science and Engineering , 2023 , 20 ( 2 ): 506 - 515 .
ZHANG Qin , ZHU Xiaoning , WANG Li . Track allocation optimization in multi-direction high-speed railway stations [J ] . Symmetry , 2019 , 11 ( 4 ): 459 .
张伯男 , 姚向明 , 赵鹏 , 等 . 分段解锁下高速铁路车站到发线运用方案调整研究 [J ] . 铁道科学与工程学报 , 2025 ( 4 ): 1432 - 1443 .
ZHANG Bonan , YAO Xiangming , ZHAO Peng , et al . Train platform scheme rescheduling in high-speed railway station with route-lock sectional-release mechanism [J ] . Journal of Railway Science and Engineering , 2025 ( 4 ): 1432 - 1443 .
张英贵 , 陈元 , 雷定猷 , 等 . 高速铁路车站到发线运用实时调整复合排序模型与算法 [J ] . 铁道学报 , 2023 , 45 ( 9 ): 34 - 45 .
ZHANG Yinggui , CHEN Yuan , LEI Dingyou , et al . Composite scheduling model and algorithm for real-time adjustment of arrival and departure track utilization in high-speed railway stations [J ] . Journal of the China Railway Society , 2023 , 45 ( 9 ): 34 - 45 .
SILVER D , HUANG A , MADDISON C J , et al . Mastering the game of Go with deep neural networks and tree search [J ] . Nature , 2016 , 529 ( 7587 ): 484 - 489 .
VINYALS O , BABUSCHKIN I , CZARNECKI W M , et al . Grandmaster level in StarCraft II using multi-agent reinforcement learning [J ] . Nature , 2019 , 575 ( 7782 ): 350 - 354 .
GUO Daya , YANG Dejian , Zhang Haowei , et al . DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning [EB/OL ] . ( 2025-01-22 )[ 2025-09-18 ] . https://arxiv.org/abs/2501.12948 https://arxiv.org/abs/2501.12948 .
LI Wenqing , NI Shaoquan . Train timetabling with the general learning environment and multi-agent deep reinforcement learning [J ] . Transportation Research Part B: Methodological , 2022 , 157 : 230 - 251 .
蒋辉 , 李博 , 贺俊源 . 基于人工智能的列车运行图智能编制技术体系框架研究 [J ] . 铁道运输与经济 , 2024 , 46 ( 1 ): 1 - 9 .
JIANG Hui , LI Bo , HE Junyuan . Research on the framework of intelligent train working diagram compilation technology system based on artificial intelligence [J ] . Railway Transport and Economy , 2024 , 46 ( 1 ): 1 - 9 .
吴卫 , 阴佳腾 , 陈照森 , 等 . 基于深度强化学习DDDQN的高速列车智能调度调整方法 [J ] . 铁道科学与工程学报 , 2024 , 21 ( 4 ): 1298 - 1308 .
WU Wei , YIN Jiateng , CHEN Zhaosen , et al . Intelligent rescheduling optimization method of high-speed railway based on deep reinforcement learning DDDQN [J ] . Journal of Railway Science and Engineering , 2024 , 21 ( 4 ): 1298 - 1308 .
范文天 , 曾勇程 , 郭一唯 , 等 . 基于强化学习的高铁列车运行图编制模型优化方法研究 [J ] . 铁道运输与经济 , 2025 , 47 ( 1 ): 70 - 81 .
FAN Wentian , ZENG Yongcheng , GUO Yiwei , et al . Optimization method of train working diagram compilation model of high speed railways based on reinforcement learning [J ] . Railway Transport and Economy , 2025 , 47 ( 1 ): 70 - 81 .
TANG Tao , CHAI Simin , WU Wei , et al . A multi-task deep reinforcement learning approach to real-time railway train rescheduling [J ] . Transportation Research Part E: Logistics and Transportation Review , 2025 , 194 : 103900 .
马驷 , 应恒 . 基于最小到达间隔的高铁车站进路分配问题研究 [J ] . 交通运输工程与信息学报 , 2020 , 18 ( 3 ): 1 - 8 .
MA Si , YING Heng . Route allocation problem of high-speed railway station based on minimum arrival interval [J ] . Journal of Transportation Engineering and Information , 2020 , 18 ( 3 ): 1 - 8 .
陈彦 , 史峰 , 秦进 , 等 . 旅客列车过站径路优化模型与算法 [J ] . 中国铁道科学 , 2010 , 31 ( 2 ): 101 - 107 .
CHEN Yan , SHI Feng , QIN Jin , et al . Optimization model and algorithm for routing passenger trains through a railway station [J ] . China Railway Science , 2010 , 31 ( 2 ): 101 - 107 .
KARP R M . Reducibility among combinatorial problems [M ] // Complexity of Computer Computations . Boston : Springer US , 1972 : 85 - 103 .
田长海 , 张守帅 , 张岳松 , 等 . 高速铁路列车追踪间隔时间研究 [J ] . 铁道学报 , 2015 , 37 ( 10 ): 1 - 6 .
TIAN Changhai , ZHANG Shoushuai , ZHANG Yuesong , et al . Study on the train headway on automatic block sections of high speed railway [J ] . Journal of the China Railway Society , 2015 , 37 ( 10 ): 1 - 6 .
王岩 , 范家铭 , 周培宇 , 等 . 基于车站进路冲突的列车间隔判定方法 : CN117644895A [P ] . 2024-03-05 .
WANG Yan , FAN Jiaming , ZHOU Peiyu , et al . A method for identifying train headway based on station route conflicts : CN117644895A [P ] . 2024-03-05 .
LUSBY R M , LARSEN J , EHRGOTT M , et al . Railway track allocation: models and methods [J ] . OR Spectrum , 2011 , 33 ( 4 ): 843 - 883 .
AHN S , SEO Y , SHIN J . Learning what to defer for maximum independent sets [C ] // Proceedings of the 37th International Conference on Machine Learning . ACM , 2020 : 134 - 144 .
DEFFERRARD M , BRESSON X , VANDERGHEYNST P . Convolutional neural networks on graphs with fast localized spectral filtering [C ] // Proceedings of the 30th International Conference on Neural Information Processing Systems . ACM , 2016 : 3844 - 3852 .
SCHULMAN J , WOLSKI F , DHARIWAL P , et al . Proximal policy optimization algorithms [EB/OL ] . ( 2017-07-20 )[ 2025-09-18 ] . https://arxiv.org/abs/1707.06347 https://arxiv.org/abs/1707.06347 .
SCHULMAN J , MORITZ P , LEVINE S , et al . High-dimensional continuous control using generalized advantage estimation [EB/OL ] . ( 2018-10-20 )[ 2025-09-18 ] . https://arxiv.org/abs/1506.02438 https://arxiv.org/abs/1506.02438 .
ANDRADE D V , RESENDE M G C , WERNECK R F . Fast local search for the maximum independent set problem [J ] . Journal of Heuristics , 2012 , 18 ( 4 ): 525 - 547 .
王宇强 , 田方晓 , 申宏楠 , 等 . 高速铁路大型客运站列车股道运用计划自动编制方法研究 [J ] . 综合运输 , 2024 , 46 ( 11 ): 68 - 74 .
WANG Yuqiang , TIAN Fangxiao , SHEN Hongnan , et al . Station track utilization for large passenger station on high speed railway [J ] . China Transportation Review , 2024 , 46 ( 11 ): 68 - 74 .
0
浏览量
5
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621