Computational & Technology Resources
an online resource for computational,
engineering & technology publications
Civil-Comp Conferences
ISSN 2753-3239
CCC: 15
PROCEEDINGS OF THE SEVENTH INTERNATIONAL CONFERENCE ON RAILWAY TECHNOLOGY: RESEARCH, DEVELOPMENT AND MAINTENANCE
Edited by: J. Pombo
Paper 14.5

From Offline Optimization to Real-Time Policy: Multi-Objective Neural Network Controllers for ATO

R. Parise1 and M. Schenker2

1Institute of Vehicle Concepts, German Aerospace Center, Berlin, Germany
2Institute of Vehicle Concepts, German Aerospace Center, Stuttgart, Germany

Full Bibliographic Reference for this paper
R. Parise, M. Schenker, "From Offline Optimization to Real-Time Policy: Multi-Objective Neural Network Controllers for ATO", in J. Pombo, (Editor), "Proceedings of the Seventh International Conference on Railway Technology: Research, Development and Maintenance ", Civil-Comp Press, Edinburgh, UK, Online volume: CCC 15, Paper 14.5, 2026, doi:10.4203/ccc.15.14.5
Keywords: automatic train operation, multi-objective optimization, neural network controller, physics-informed neural network, differentiable physics simulation, real-time control, optimal control.

Abstract
The transition towards autonomous trains necessitates intelligent automatic train operation systems capable of real-time, multi-objective trajectory optimization for highly efficient operations. While traditional open-loop optimal control methods are computationally demanding for online adaptation, data-driven approaches such as behavioral cloning inherently suffer from compounding errors, and standard reinforcement learning struggles to satisfy safety constraints in sparse-reward settings. To bridge this gap, this paper presents a novel learning framework for robust, closed-loop train control. We propose a two-step methodology: first, a neural policy is warm-started through pre-training on a dataset of constraint-satisfying open-loop trajectories generated via numerical optimization. Subsequently, the policy is fine-tuned in closed-loop via direct policy search within a continuous neural ordinary differential equation architecture. By backpropagating gradients directly through the system dynamics, a curriculum learning strategy is used to optimize multi-objective costs such as energy consumption, wear and punctuality. Simulation results demonstrate that this physics-informed approach effectively eliminates the compounding errors, adheres to nonlinear speed boundaries, and achieves real-time operational robustness. The resulting controller evaluates in milliseconds and synthesizes energy-efficient driving profiles under varying operational disturbances. This framework enables automated trains to adapt their driving style on the tracks, reducing energy consumption and ensuring a more comfortable, on-time journey.

download the full-text of this paper (PDF, 16 pages, 570 Kb)

go to the previous paper
go to the next paper
return to the table of contents
return to the volume description