Multimodal Diagnostics for Wilson’s Disease Bridging Object Detection, Large Language Model, and Speech Synthesis
摘要
This research presents an innovative multimodal approach for diagnosing Wilson’s disease by integrating linguistic information with object detection methods, thereby enhancing interpretability and accuracy in MRI-based diagnostics. Utilizing deep learning models, specifically transformers for object detection, we incorporate language-driven analysis to detect disease-specific biomarkers more effectively in brain MRI scans. This approach leverages linguistic cues from medical images, allowing the model to better align visual information with diagnostic terms relevant to Wilson’s disease. Additionally, by incorporating voice synthesis alongside visual and textual modalities, our framework offers a richer, multimodal diagnostic tool that improves model confidence and accuracy, surpassing single-modality detection methods. This comprehensive setup underscores the potential of multimodal AI in supporting accurate and context-aware clinical decision-making.