Forecasting Government Project Costs in Colombia: Combining Regression-Based and Text-Mining Approaches for Predictive Analysis
摘要
Understanding the projected costs of projects within various sectors of a country is crucial for resource allocation and timely delivery. In Colombia, comprehensive government project data is accessible through the National Government’s open data platform. Utilizing these datasets from the National Planning Department, we construct a predictive model leveraging regression analysis to estimate the expenses associated with governmental initiatives. This work evaluates several regression models, using diverse evaluation error metrics, to determine the most effective model for deployment. A key component of our approach is to combine textual attributes into a single variable, and subsequently apply text mining techniques, in order to obtain insights from free text fields in the data sets. Ultimately, the Adaboost model combined with TF-IDF emerged as the most precise combination of models, exhibiting a mean average precision error (MAPE) of 17.6%, closely followed by the Random Forest model combined with TF-IDF with a MAPE of 17.9%.