The photovoltaic (PV) power forecast is a key aspect of grid stability and optimal allocation of energy resources. However, accurate prediction is impaired due to variability in PV production from meteorological fluctuations. This study presents the performance of two sophisticated deep learning models, LSTM and CNN-LSTM, that have been trained on datasets generated using two methods: correlation coefficients (Important_data) and Principal Component Analysis (PCA_data) preprocessing techniques. The analysis was performed on two temporal horizons: one-step forecasting (5 min), and six-step forecasting (30 min). As seen in the results, Important_data based on correlation coefficients always provides stronger predictive consistencies than PCA_data. For example, for one step (5 min), when comparing MAE, RMSE, and R2, the LSTM model on Important_data achieves MAE = 0.015, RMSE = 0.044, and R2 = 0.9810, giving better performance than those on PCA_data (MAE = 0.018, RMSE = 0.047, and R2 = 0.9780). The results indicate that preserving essential features during the preprocessing of small PV datasets can be done more effectively using a correlation coefficient-based technique due to its lower loss of information.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Feature Selection Techniques for Accurate Photovoltaic Power Forecasting: A Study of Correlation Coefficients Versus PCA

  • Samira Marhraoui,
  • Hassan Silkan,
  • Said Laasri

摘要

The photovoltaic (PV) power forecast is a key aspect of grid stability and optimal allocation of energy resources. However, accurate prediction is impaired due to variability in PV production from meteorological fluctuations. This study presents the performance of two sophisticated deep learning models, LSTM and CNN-LSTM, that have been trained on datasets generated using two methods: correlation coefficients (Important_data) and Principal Component Analysis (PCA_data) preprocessing techniques. The analysis was performed on two temporal horizons: one-step forecasting (5 min), and six-step forecasting (30 min). As seen in the results, Important_data based on correlation coefficients always provides stronger predictive consistencies than PCA_data. For example, for one step (5 min), when comparing MAE, RMSE, and R2, the LSTM model on Important_data achieves MAE = 0.015, RMSE = 0.044, and R2 = 0.9810, giving better performance than those on PCA_data (MAE = 0.018, RMSE = 0.047, and R2 = 0.9780). The results indicate that preserving essential features during the preprocessing of small PV datasets can be done more effectively using a correlation coefficient-based technique due to its lower loss of information.