<p>Corporate revenue prediction is essential for strategic planning and decision-making. However, challenges arise when data are limited in terms of sample size, distribution and feature availability. This study explores machine learning (ML)-based corporate revenue prediction under data-constrained scenarios. Four ML models—Support Vector Machine (SVM), Artificial Neural Network (ANN), K-Nearest Neighbors (KNN), and Random Forest (RF)—were developed and optimized via Bayesian Optimization (BO) to enhance predictive performance. Among them, the BO-RF model achieved the best results with an accuracy of 0.8940 and a macro AUC of 0.9745. SHAP analysis identified “total assets,” “wages payable,” and “employee count” as key features. The study demonstrates that even under limited data conditions, reliable category prediction is achievable through proper model selection and hyperparameter tuning. These findings offer practical guidance for revenue assessment in data-constrained environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Corporate revenue category prediction with limited data through parameter optimization and model comparison

  • Mingyang Zhang,
  • Hao Wang,
  • Qinglin Meng,
  • Yuwei Zhai,
  • Jian Zuo,
  • Na Dong

摘要

Corporate revenue prediction is essential for strategic planning and decision-making. However, challenges arise when data are limited in terms of sample size, distribution and feature availability. This study explores machine learning (ML)-based corporate revenue prediction under data-constrained scenarios. Four ML models—Support Vector Machine (SVM), Artificial Neural Network (ANN), K-Nearest Neighbors (KNN), and Random Forest (RF)—were developed and optimized via Bayesian Optimization (BO) to enhance predictive performance. Among them, the BO-RF model achieved the best results with an accuracy of 0.8940 and a macro AUC of 0.9745. SHAP analysis identified “total assets,” “wages payable,” and “employee count” as key features. The study demonstrates that even under limited data conditions, reliable category prediction is achievable through proper model selection and hyperparameter tuning. These findings offer practical guidance for revenue assessment in data-constrained environments.