<p>This paper develops a hybrid framework for quantifying the financial impact of data breaches by combining predictive machine learning with extreme value theory (EVT). Using incident-level breach data from the Privacy Rights Clearinghouse (PRC) covering the period 2005–2020, we first estimate the number of compromised records with a Random Forest model trained on organizational, temporal and attack-type characteristics. We then analyze the tail behavior of the predicted losses to capture the fat-tailed distribution of cyber risks. Our results indicate that the distribution of affected records is well represented by a Fréchet law, and we estimate the parameters of the Generalized Extreme Value (GEV) distribution to compute Value-at-Risk (VaR) at high confidence levels. This two-stage approach provides a rigorous assessment of maximum potential losses, addressing the question of cyber-risk insurability. By linking predictive accuracy with tail risk quantification, our findings deliver actionable insights for insurers, regulators, and organizations seeking to anticipate and manage the financial consequences of large-scale data breaches.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CyberRisk Prediction using Machine Learning and Extreme Value Theory

  • Jules SADEFO KAMDEM,
  • Danielle Selambi Kapsa

摘要

This paper develops a hybrid framework for quantifying the financial impact of data breaches by combining predictive machine learning with extreme value theory (EVT). Using incident-level breach data from the Privacy Rights Clearinghouse (PRC) covering the period 2005–2020, we first estimate the number of compromised records with a Random Forest model trained on organizational, temporal and attack-type characteristics. We then analyze the tail behavior of the predicted losses to capture the fat-tailed distribution of cyber risks. Our results indicate that the distribution of affected records is well represented by a Fréchet law, and we estimate the parameters of the Generalized Extreme Value (GEV) distribution to compute Value-at-Risk (VaR) at high confidence levels. This two-stage approach provides a rigorous assessment of maximum potential losses, addressing the question of cyber-risk insurability. By linking predictive accuracy with tail risk quantification, our findings deliver actionable insights for insurers, regulators, and organizations seeking to anticipate and manage the financial consequences of large-scale data breaches.