Two-set Variable Framework for Precise Vehicle Speed Prediction
摘要
Accurate speed prediction of vehicles is significant in traffic safety, traffic flow control, and environmentally friendly transport planning. Most existing research addresses these variables separately or utilizes neural networks and traditional statistical models, with limited use of interpretable machine learning methods and inadequate attention to multicollinearity, sensitivity testing, vehicle-type heterogeneity, and dataset robustness. To fill these gaps, this study presents a two-variable set framework for analyzing vehicle-specific and environmental variables separately and in combination, utilizing RR, DTR, LASSO, and their combined models. The structure comprises Monte Carlo-based data augmentation, multicollinearity checks via the VIF, K-means clustering for classifying cars by type, the FAST for assessing variable influence, and Wilcoxon rank-sum tests for statistical verification. The dataset, collected by NTNU from a rural two-lane highway (March 2012–April 2014, 135,928 observations with approximately 10% heavy-duty vehicles), is used to evaluate model performance under different data conditions. In the case of the Monte Carlo-generated datasets, the hybrid models, especially DTEE, were able to produce R² that were close to 0.99. This is an indication of the behavior of the models in terms of the smoothed data, which can be considered an upper bound estimate. Thus, the results of cross-validation obtained based on the observational data constitute the major evidence for practical predictive ability, while the Monte Carlo simulation serves as a sensitivity test showing how the framework behaves in low-noise environment.