Abstract
Accurate crop yield forecasting requires identifying informative predictors from high-dimensional meteorological time series with multiple time lags. Such settings often lead to overfitting, reduced interpretability, and increased computational cost. This paper proposes a two-phase feature-selection framework, termed CS \(_\text {MRMR}\) -DWES \(_R\) , to address these challenges in agricultural time-series regression. In the first phase, a supervised filter ranks features using correlation statistics with a maximum relevance-minimum redundancy criterion, retaining a compact subset of informative variables. In the second phase, a wrapper-based discrete weighted evolutionary strategy automatically selects the optimal feature subset by directly optimizing the predictive performance of a SARIMAX regression model. The proposed framework is evaluated on 20 years of monthly avocado and mango yield data from the Axarquía region of Málaga in Spain under multiple lag configurations. This feature selection might slightly change over time and depends entirely on the crop type and geographical location. The proposed CS \(_\text {MRMR}\) -DWES \(_\text {R}\) framework inherently requires high-performance computing (HPC) capabilities due to the scale and complexity of the agricultural time-series datasets under study. Experimental results show that CS \(_\text {MRMR}\) -DWES \(_R\) improves forecasting accuracy while reducing feature dimensionality and computational cost compared with baseline and existing feature-selection methods. These findings highlight the effectiveness of the proposed approach for high-dimensional agricultural time-series analysis.
Graphic abstract