<p>The F1 score is a widely used metric for evaluating classification models, particularly in imbalanced data scenarios. However, conventional methods typically provide only a single point estimate, which fails to capture the uncertainty in estimating the F1 score from finite samples, uncertainty that is exacerbated when few minority-class outcomes are observed. In this study, we propose a practical Bayesian framework for estimating the posterior distribution of the F1 score using the Dirichlet-Multinomial model with sequential updates. This approach enables continuous and interpretable uncertainty quantification. It is particularly well suited for streaming data and online learning environments, where model evaluation must be updated as new observations arrive. We validate the proposed method through extensive simulation studies that examine its performance across various levels of class imbalance and sample sizes. In addition, we demonstrate its practical utility through application experiments on benchmark classification tasks, highlighting how our framework provides interpretable uncertainty measures and enhances real-time model monitoring. Overall, the method supports more robust and informed decision-making by offering a dynamic understanding of model performance over time.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sequential Bayesian estimation of the F1 score using the Dirichlet-multinomial model

  • Surani Matharaarachchi,
  • Maxime Turgeon,
  • Mike Domaratzki,
  • Saman Muthukumarana

摘要

The F1 score is a widely used metric for evaluating classification models, particularly in imbalanced data scenarios. However, conventional methods typically provide only a single point estimate, which fails to capture the uncertainty in estimating the F1 score from finite samples, uncertainty that is exacerbated when few minority-class outcomes are observed. In this study, we propose a practical Bayesian framework for estimating the posterior distribution of the F1 score using the Dirichlet-Multinomial model with sequential updates. This approach enables continuous and interpretable uncertainty quantification. It is particularly well suited for streaming data and online learning environments, where model evaluation must be updated as new observations arrive. We validate the proposed method through extensive simulation studies that examine its performance across various levels of class imbalance and sample sizes. In addition, we demonstrate its practical utility through application experiments on benchmark classification tasks, highlighting how our framework provides interpretable uncertainty measures and enhances real-time model monitoring. Overall, the method supports more robust and informed decision-making by offering a dynamic understanding of model performance over time.