Meta-Learning-Based Imputation in Solar Energy Systems: Bridging the Gap Between Missing Data and Forecasting Reliability
摘要
Solar energy forecasting relies on continuous, high-quality data, but sensor faults, communication failures, and harsh weather often lead to missing values, reducing reliability. This study introduces a meta-learning-based imputation framework to address such data gaps in solar energy systems. The proposed Stacked Ensemble Imputation (SEI) model integrates three base estimators—Time Series k-Nearest Neighbors (TS-kNN), Adaptive Rolling Mean, and Linear Interpolation—whose outputs are fused using a Random Forest meta-learner. Recursive Feature Elimination (RFE) is employed to optimize feature selection, enhancing model interpretability and reducing computational load. The model was evaluated on a 3 MW solar dataset with artificially induced missing values ranging from 10% to 50%. SEI consistently achieved superior performance, with a minimum Mean Absolute Error of 0.041 and R2 values exceeding 0.99 across all scenarios. These results demonstrate the model’s ability to preserve temporal continuity and original data structures even under severe data loss. The SEI framework offers a scalable and efficient solution for solar data preprocessing, significantly improving forecasting accuracy. Future work will focus on real-time deployment and extending the approach to other renewable energy domains.