Interpretable Machine Learning for Public Bus Service Efficiency: A SHAP-Driven Framework for Operational Analytics
摘要
Public bus transport systems in emerging economies face persistent challenges in achieving operational efficiency, service coverage, and financial sustainability. This study presents an integrated analytical framework that combines predictive modeling, explainable artificial intelligence (XAI), and unsupervised clustering to diagnose and optimize fleet operations, with application to the Andhra Pradesh State Road Transport Corporation (APSRTC) in India. High-resolution planned and monitored operational datasets were used to train Light Gradient Boosting Machine (LightGBM), Random Forest, and Ridge Regression models for predicting fuel consumption. Model interpretability was achieved through SHapley Additive exPlanations (SHAP) to measure the relative contribution of predictor variables, while k-means clustering in a principal component analysis (PCA) reduced feature space was applied to identify service groups with distinct operational characteristics. The best performing models produced root mean squared error (RMSE) values between 18.18 and 20.48 L and coefficient of determination (R2) values up to 0.71. Trip distance consistently appeared as the most influential predictor in both planned and monitored datasets. Clustering analysis identified three operational groups: long-haul high consumption, short-haul low consumption, and intermediate services, providing a basis for targeted operational improvement. The integration of predictive modeling, SHAP-based interpretability, and operational segmentation enables two levels of insight: fleet-level performance diagnosis and route-specific strategic analysis. The findings support data-driven recommendations for improving fuel efficiency, reducing operational costs, and optimizing service planning. The proposed framework offers a replicable approach for state-owned transport corporations operating in similar contexts.