错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal random subspace for breast cancer molecular subtypes prediction by integrating multi-dimensional data

  • Fatima-Zahrae Nakach,
  • Ali Idri,
  • Gbègninougbo Aurel Davy Tchokponhoue

摘要

Accurate identification of breast cancer molecular subtypes significantly impacts patient prognosis and treatment decisions. Multimodal fusion techniques have shown promise in improving the performance of deep learning models by leveraging the complementary strengths of various modalities. However, the integration of multi-dimensional data poses a challenge due to the heterogeneity and complexity of the data sources. Random subspace can improve the robustness and generalization abilities of multimodal classification techniques by reducing variance and making it easier to identify the most informative features. The present study suggests a new approach referred to as Multimodal Random Subspace Support Vector Machine (MRSVM) ensemble that effectively combines multidimensional data encompassing copy number variation, clinical information, gene expression data and histopathological whole slide images to improve the prediction of breast cancer molecular subtypes. To highlight the prowess and distinctive qualities of the newly devised approach, we conducted an extensive comparative examination across all possible combinations of these four modalities, where we were able to address six research questions, which mainly involve identifying the optimal combination of the four available modalities, the most effective fusion strategy, and the most suitable classification technique. Results indicated that early fusion models outperformed late fusion models, notably, the MRSVM ensemble based on early fusion achieved the highest accuracy of 88.07% for molecular subtype prediction over the subtypes (Normal-like, Luminal A, Luminal B, HER2-enriched, and Basal-like). The findings also confirmed that multimodal models performed better than classifiers built on an individual data modality, and therefore allows us to suggest their use in place of monomodal models when multiple data modalities are available especially with ensemble methods.