<p>This paper presents a reproducible benchmark for irradiance-driven PV power prediction using five years of hourly NASA POWER data (2020–2024) for Hengsha Island, Shanghai. Because measured plant output was unavailable, the response variable is an engineered normalized PV proxy derived from irradiance and temperature, so the reported results should be interpreted as benchmark evidence rather than field validation against measured SCADA power. The workflow applies leakage-safe chronological partitioning, common feature engineering, daylight-only evaluation, and persistence-style skill comparison across XGBoost, Random Forest, ANFIS-SC, GRU, LSTM, and CNN-BiGRU-AM. On the 2024 test set, tree ensembles achieve the strongest benchmark accuracy (XGBoost: <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(R^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>R</mi> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> = 0.9994, RMSE = 0.0009; Random Forest: <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(R^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>R</mi> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> = 0.9978, RMSE = 0.0018), ANFIS-SC offers an interpretable alternative (<InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(R^2\)</EquationSource> <EquationSource Format="MATHML"><math> <msup> <mi>R</mi> <mn>2</mn> </msup> </math></EquationSource> </InlineEquation> = 0.9886), recurrent networks show moderate performance with higher computational cost, and the tested CNN-BiGRU-AM configuration is unstable in this setting (negative skill). Additional ablation and parameter-sensitivity analyses clarify that this negative result is specific to the tested attention-enhanced hybrid and that the benchmark ranking is robust to plausible changes in the engineered PV-target parameters. The benchmark clarifies trade-offs among accuracy, interpretability, and complexity, and the accompanying open-source implementation supports reproducible comparison and further methodological extension.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Intelligent and Interpretable Models for Irradiance-Driven PV Power Prediction

  • Muhammad Abbas,
  • Bulut Hüner,
  • Duanjin Zhang

摘要

This paper presents a reproducible benchmark for irradiance-driven PV power prediction using five years of hourly NASA POWER data (2020–2024) for Hengsha Island, Shanghai. Because measured plant output was unavailable, the response variable is an engineered normalized PV proxy derived from irradiance and temperature, so the reported results should be interpreted as benchmark evidence rather than field validation against measured SCADA power. The workflow applies leakage-safe chronological partitioning, common feature engineering, daylight-only evaluation, and persistence-style skill comparison across XGBoost, Random Forest, ANFIS-SC, GRU, LSTM, and CNN-BiGRU-AM. On the 2024 test set, tree ensembles achieve the strongest benchmark accuracy (XGBoost: \(R^2\) R 2 = 0.9994, RMSE = 0.0009; Random Forest: \(R^2\) R 2 = 0.9978, RMSE = 0.0018), ANFIS-SC offers an interpretable alternative ( \(R^2\) R 2 = 0.9886), recurrent networks show moderate performance with higher computational cost, and the tested CNN-BiGRU-AM configuration is unstable in this setting (negative skill). Additional ablation and parameter-sensitivity analyses clarify that this negative result is specific to the tested attention-enhanced hybrid and that the benchmark ranking is robust to plausible changes in the engineered PV-target parameters. The benchmark clarifies trade-offs among accuracy, interpretability, and complexity, and the accompanying open-source implementation supports reproducible comparison and further methodological extension.