Computational & Technology Resources
an online resource for computational,
engineering & technology publications
Civil-Comp Conferences
ISSN 2753-3239
CCC: 15
PROCEEDINGS OF THE SEVENTH INTERNATIONAL CONFERENCE ON RAILWAY TECHNOLOGY: RESEARCH, DEVELOPMENT AND MAINTENANCE
Edited by: J. Pombo
Paper 21.3

Feasibility of Vision–Language Model–Based Collision Risk Assessment as a Driver Assistance System for Tram Drivers: A Preliminary Study

S. Hovanotayan1, W. Wang1, K. Nakano1, T. Takata2, M. Ueda3, K. Sakakidani3 and T. Maeoka3

1Institute of Industrial Science, The University of Tokyo, Japan
2, Kyosan Electric Manufacturing Co., Ltd., Yokohama, Japan
3, Hiroshima Electric Railway Co., Ltd., Japan

Full Bibliographic Reference for this paper
S. Hovanotayan, W. Wang, K. Nakano, T. Takata, M. Ueda, K. Sakakidani, T. Maeoka, "Feasibility of Vision–Language Model–Based Collision Risk Assessment as a Driver Assistance System for Tram Drivers: A Preliminary Study", in J. Pombo, (Editor), "Proceedings of the Seventh International Conference on Railway Technology: Research, Development and Maintenance ", Civil-Comp Press, Edinburgh, UK, Online volume: CCC 15, Paper 21.3, 2026, doi:10.4203/ccc.15.21.3
Keywords: tram, safety, collision risk assessment, artificial intelligence, computer vision, vision-language model, driver assistance system.

Abstract
As trams typically share traffic spaces with cars, cyclists and pedestrians, collisions between trams and surrounding road agents have become an increasingly critical problem. Assessing potential risk is undeniably an effective measure to mitigate such safety issues. Nevertheless, comprehensively grasping tramway scenarios remains challenging due to their inherent complexity compared to other traffic environments. Recently, vision-language models (VLMs) have emerged as promising tools for visual understanding in various fields, especially in autonomous driving, because of their advanced scene analysis and reasoning abilities. In this paper, we investigate the feasibility of deploying VLMs for collision risk assessment dedicated to tram contexts. Existing VLMs, including GPT-5, Gemini 3 Pro, Sonnet 4.5, Llama 3.2 Vision, Qwen 3 VL and DeepSeek VL 2, are selected. Firstly, a dataset of tram scenarios is collected from a tram driving simulator, ranging from safe to dangerous cases. Each VLM is prompted to analyse scene features, assess collision risk and suggest driving actions in the given scenarios. In the preliminary tests, each model successfully managed to perform its designated tasks with promising performance. In future work, the VLMs’ outputs will be compared with evaluations by professional tram drivers obtained through questionnaire surveys, thereby serving as ground truth data.

download the full-text of this paper (PDF, 14 pages, 1387 Kb)

go to the previous paper
go to the next paper
return to the table of contents
return to the volume description