错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Encoding clinical preferences via resampling: an application to pain assessment

  • Miguel Carvalho,
  • Daniela Pais,
  • Raquel Sebastião,
  • Armando Pinho,
  • Susana Brás

摘要

Class imbalance can significantly impact the performance of learning algorithms, often leading to prediction bias toward the majority class. This challenge is particularly critical in healthcare-related domains, as medical datasets are often imbalanced, hindering the accurate prediction of the minority class, which is commonly the class of interest. As such, this work introduces a novel resampling algorithm, designated Genetic Beta Oversampling, which integrates user-defined preferences into the synthetic data generation process, allowing fine control over the model’s inclination towards false negatives or false positives. These user preferences are encoded in the form of a parameter, \(\beta \) , which dictates the trade-off between recall and precision that the method should seek to achieve. This flexibility is particularly relevant in clinical settings, where prioritizing recall can enhance patient care by reducing missed diagnoses. We evaluate the proposed approach on the EMPA and AI4PAIN datasets for pain classification, a recall-critical task in which undetected pain episodes must be minimized. Experimental results show that our method consistently surpasses SMOTE, SMOTE-IPF, and four cost-sensitive classifiers in terms of F \(\beta \) -score across a range of \(\beta \) values and synthetically induced imbalance ratios. These results highlight the adaptability of the method to recall-sensitive applications, including pain assessment and broader clinical decision-support scenarios.