A Multimodal Deep Learning Framework for Mycetoma Classification: Integrating Vision Transformers, Medical Language Models, and Transfer Learning
摘要
Mycetoma is a chronic neglected tropical disease that progressively destroys skin, subcutaneous tissue, and bone. Differentiating its bacterial (Actinomycetoma) and fungal (Eumycetoma) forms is vital for correct treatment but remains difficult due to overlapping histopathological features and limited data. This study introduces a two-phase deep learning framework for automated Mycetoma classification from histopathological images. In Phase I, optimized InceptionV3 fine-tuning with augmented data (2052 images) achieved 94.66% accuracy (AUC 0.983). Phase II evaluated transformer and multimodal models, including a vision-only MedicalMultimodalLLM (DeiT-Base) and a TRUE Multimodal Medical LLM combining DeiT with PubMedBERT text embeddings. These achieved accuracies of 99.51% and near-perfect accuracy ( \(\approx \) 100%) and recall (FN = 0 across test splits), respectively. Results demonstrate the synergy of optimized transfer learning and vision-language modeling for small medical datasets. The proposed framework establishes a reproducible benchmark for Mycetoma histopathology and represents a step toward AI-assisted diagnosis of neglected tropical diseases.