<p>Machine learning (ML) models are being extensively used in existing research efforts to diagnose diseases in a better way and offer solutions to medical problems. However, it is necessary to focus on bias or disparity concerning the algorithm and demographic attributes, like gender, in automatic disease detection (ADD). Even though there are some research efforts to consider or mitigate the several disparities of digital ML-based systems, the number of studies is still limited, and existing studies focus only on a single specific disease type, mostly in a different manner. Hence, this study performs an extensive evaluation of eight different ML pipelines to inspect the superiority of gender-sensitive, stacking-ensemble, and the hybrid of the previous two approaches. To the best of our knowledge, this is the first study in the literature that handles ML-based ADD with a broad perspective that covers different disease types, unlike its peers. Extensive evaluations performed on six datasets of four disease types show that the performances measured by f1-score vary in a range between 0.743 and 0.965 depending on the nature of the datasets. On the other hand, a gender-sensitive approach mitigates the performance disparity in only a few cases and does not improve the overall prediction performance in the majority of cases. Contrarily, the stacking-ensemble approach is found to be superior in almost all cases, considering the best overall performance. The behavior of pipelines involving these approaches also shows similar behavior regardless of the size and balance status of the dataset at hand.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Addressing Gender Disparities in Automatic Disease Detection: A Broad-Spectrum Analysis of Machine Learning Across Diverse Disease Types

  • Önder Çoban,
  • Ayşe Kartal

摘要

Machine learning (ML) models are being extensively used in existing research efforts to diagnose diseases in a better way and offer solutions to medical problems. However, it is necessary to focus on bias or disparity concerning the algorithm and demographic attributes, like gender, in automatic disease detection (ADD). Even though there are some research efforts to consider or mitigate the several disparities of digital ML-based systems, the number of studies is still limited, and existing studies focus only on a single specific disease type, mostly in a different manner. Hence, this study performs an extensive evaluation of eight different ML pipelines to inspect the superiority of gender-sensitive, stacking-ensemble, and the hybrid of the previous two approaches. To the best of our knowledge, this is the first study in the literature that handles ML-based ADD with a broad perspective that covers different disease types, unlike its peers. Extensive evaluations performed on six datasets of four disease types show that the performances measured by f1-score vary in a range between 0.743 and 0.965 depending on the nature of the datasets. On the other hand, a gender-sensitive approach mitigates the performance disparity in only a few cases and does not improve the overall prediction performance in the majority of cases. Contrarily, the stacking-ensemble approach is found to be superior in almost all cases, considering the best overall performance. The behavior of pipelines involving these approaches also shows similar behavior regardless of the size and balance status of the dataset at hand.