<p>Predictive maintenance for urban electric transport operates under severe class imbalance: failures are rare but individually costly. We benchmark thirty classical, boosting, and deep tabular models on an anonymised multi-modal fleet dataset (trams, trolleybuses, electric buses) under a leakage-resilient <InlineEquation ID="IEq1"><EquationSource Format="TEX">\(5\times 5\)</EquationSource></InlineEquation> repeated stratified cross-validation protocol, and evaluate seven single-objective and three multi-objective hyperparameter-optimisation strategies for a reference gradient-boosting model. Re-evaluating every tuned configuration under the same protocol used to rank the models, we find that the gains reported by the optimisers are artefacts of cross-validation optimism: the inner objective that the search maximises is inflated in this rare-event regime (mean optimism <InlineEquation ID="IEq2"><EquationSource Format="TEX">\(+0.17\)</EquationSource></InlineEquation> ROC-AUC), its strategy ranking does not survive matched evaluation, and the degree of over-fitting grows monotonically with how aggressively a strategy maximises the inner score. Under matched evaluation no tuned configuration outperforms a sensible default. An operational, cost-sensitive analysis shows that, at the prevailing failure rate, even the best model has no usable decision threshold. A feature-group ablation locates the residual signal in energy and kinematics and in service-cycle history, and shows it saturates at daily telemetry resolution. We release the benchmark, code, and an honest evaluation protocol for rare-event predictive maintenance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-validation optimism undermines hyperparameter optimisation in rare-event predictive maintenance: a leakage-resilient benchmark for urban electric transport

  • Nikita V. Martyushev,
  • Boris V. Malozyomov,
  • Sergei O. Kurashkin,
  • Ivan P. Malashin,
  • Vadim S. Tynchenko,
  • Aleksei S. Borodulin,
  • Ahmad Hammoud,
  • Connie Tee

摘要

Predictive maintenance for urban electric transport operates under severe class imbalance: failures are rare but individually costly. We benchmark thirty classical, boosting, and deep tabular models on an anonymised multi-modal fleet dataset (trams, trolleybuses, electric buses) under a leakage-resilient \(5\times 5\) repeated stratified cross-validation protocol, and evaluate seven single-objective and three multi-objective hyperparameter-optimisation strategies for a reference gradient-boosting model. Re-evaluating every tuned configuration under the same protocol used to rank the models, we find that the gains reported by the optimisers are artefacts of cross-validation optimism: the inner objective that the search maximises is inflated in this rare-event regime (mean optimism \(+0.17\) ROC-AUC), its strategy ranking does not survive matched evaluation, and the degree of over-fitting grows monotonically with how aggressively a strategy maximises the inner score. Under matched evaluation no tuned configuration outperforms a sensible default. An operational, cost-sensitive analysis shows that, at the prevailing failure rate, even the best model has no usable decision threshold. A feature-group ablation locates the residual signal in energy and kinematics and in service-cycle history, and shows it saturates at daily telemetry resolution. We release the benchmark, code, and an honest evaluation protocol for rare-event predictive maintenance.