In disease classification, multimodal fusion model has emerged as a promising approach to enhance both diagnostic accuracy and efficiency. In this chapter, we propose to train the multimodal model for classifying respiratory illnesses from chest X-ray images and symptom texts. The proposed model uses Support Vector Machine (SVM) on top of Vision Transformer (ViT) and Random Forest (RF) with Term Frequency-Inverse Document Frequency (TF-IDF) for the disease classification. We collect a new real dataset from An Giang province regional general hospital. The experimental results show that using only the ViT model for classifying diseases from chest X-ray images achieves an accuracy of 62.43%. Furthermore, disease classification based on clinical symptom text achieves accuracies of 77.52% and 70.10% respectively when using RF and SVM models with TF-IDF technique. Our multimodal model achieves an accuracy with 80.10%, surpassing unimodal models using only images or text.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Image–Text Fusion for Respiratory Disease Classification

  • Thi-Diem Truong,
  • Phuoc-Hai Huynh,
  • Van Hoa Nguyen,
  • Thanh-Nghi Do

摘要

In disease classification, multimodal fusion model has emerged as a promising approach to enhance both diagnostic accuracy and efficiency. In this chapter, we propose to train the multimodal model for classifying respiratory illnesses from chest X-ray images and symptom texts. The proposed model uses Support Vector Machine (SVM) on top of Vision Transformer (ViT) and Random Forest (RF) with Term Frequency-Inverse Document Frequency (TF-IDF) for the disease classification. We collect a new real dataset from An Giang province regional general hospital. The experimental results show that using only the ViT model for classifying diseases from chest X-ray images achieves an accuracy of 62.43%. Furthermore, disease classification based on clinical symptom text achieves accuracies of 77.52% and 70.10% respectively when using RF and SVM models with TF-IDF technique. Our multimodal model achieves an accuracy with 80.10%, surpassing unimodal models using only images or text.