In this paper, we evaluate bagging ensembles for breast cancer diagnosis using a multimodal dataset “PathoEMR” which combines pathology images and electronic medical records. We implement the bagging ensemble method to two deep learning models (Densenet201 and VGG16) for pathological images and seven machine learning classifiers for electronic medical records. The best bagging ensemble for electronic medical records achieved an accuracy mean value of 92.31%, a precision of 95.56%, a recall of 90.66%, an F1-score of 92.99%, and a specificity of 95.87%. While the best bagging ensemble for the classification of pathological images achieved an accuracy mean value of 85.59%, a precision of 86.87%, a recall of 90.45%, and an f1-score of 88.60%. We compare the bagging ensembles with different number of base learners between them and with the single models based on the Scott-Knott statistical method and the Borda count voting method. All our bagging ensemble models outperform the single models for breast cancer classification. Moreover, we found that the best bagging ensembles outperform the state-of-the-art models applied to the same dataset, which achieved an accuracy of 83.60% with VGG16 applied to pathological images and 78.50% with the auto-encoder applied to electronic medical records.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of Bagging Ensembles on Multimodal Data for Breast Cancer Diagnosis

  • Abdulganiyu Jimoh,
  • Fatima-Zahrae Nakach,
  • Ali Idri,
  • Ikram Chairi

摘要

In this paper, we evaluate bagging ensembles for breast cancer diagnosis using a multimodal dataset “PathoEMR” which combines pathology images and electronic medical records. We implement the bagging ensemble method to two deep learning models (Densenet201 and VGG16) for pathological images and seven machine learning classifiers for electronic medical records. The best bagging ensemble for electronic medical records achieved an accuracy mean value of 92.31%, a precision of 95.56%, a recall of 90.66%, an F1-score of 92.99%, and a specificity of 95.87%. While the best bagging ensemble for the classification of pathological images achieved an accuracy mean value of 85.59%, a precision of 86.87%, a recall of 90.45%, and an f1-score of 88.60%. We compare the bagging ensembles with different number of base learners between them and with the single models based on the Scott-Knott statistical method and the Borda count voting method. All our bagging ensemble models outperform the single models for breast cancer classification. Moreover, we found that the best bagging ensembles outperform the state-of-the-art models applied to the same dataset, which achieved an accuracy of 83.60% with VGG16 applied to pathological images and 78.50% with the auto-encoder applied to electronic medical records.