<p>Evaluation metrics provide a means for quantifying and comparing performances of supervised learning models, but drawing meaningful conclusions from acquired scores requires a contextual framework. Our paper addresses this by introducing the Dutch scaler (DS), a novel performance indicator for binary classification models. It quantifies a model’s learning by contextualizing empirical metric scores with a baseline (Dutch draw) and a new instrument (Dutch oracle) representing the prediction quality of an “optimal” classifier. The DS performance indicator expresses the relative contribution of these components to obtain a model’s score, specifying the actual learning quality. We derived closed-form expressions to map metric scores to DS scores for common evaluation metrics and categorized them by their functional form and second derivative. The DS enhances the assessment of classifiers and facilitates a framework to compare prediction quality differences between models with varying metric scores.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Dutch Scaler Performance Indicator: How Much Did My Model Actually Learn?

  • Etienne Pieter van de Bijl,
  • Jan Gerard Klein,
  • Joris Pries,
  • Sandjai Bhulai,
  • Robert Douwe van der Mei

摘要

Evaluation metrics provide a means for quantifying and comparing performances of supervised learning models, but drawing meaningful conclusions from acquired scores requires a contextual framework. Our paper addresses this by introducing the Dutch scaler (DS), a novel performance indicator for binary classification models. It quantifies a model’s learning by contextualizing empirical metric scores with a baseline (Dutch draw) and a new instrument (Dutch oracle) representing the prediction quality of an “optimal” classifier. The DS performance indicator expresses the relative contribution of these components to obtain a model’s score, specifying the actual learning quality. We derived closed-form expressions to map metric scores to DS scores for common evaluation metrics and categorized them by their functional form and second derivative. The DS enhances the assessment of classifiers and facilitates a framework to compare prediction quality differences between models with varying metric scores.