Enhancing Diabetic Foot Ulcer Assessment Through Fine-Tuned Vision-Language Models
摘要
Diabetic foot ulcers (DFUs) are a significant cause of lower limb amputations and hospitalizations, placing a substantial burden on patients and healthcare systems. Early assessments are critical for preventing complications, yet the shortage of specialists cannot meet the widespread prevalence of this condition. This paper explores augmenting the DFU clinical workflow by applying vision-language models for generating clinically relevant assessments of DFUs from images. Using the Wound-Ischemia-Foot Infection (WIfI) classification system as a structured framework, we assess the performance of LLaVA-Mistral models fine-tuned on annotated DFU datasets compared to LLaVA-Mistral and GPT-4o baselines. Our findings demonstrate that the initial fine-tuned LLaVA-Mistral model achieved on average 16% higher average accuracy in predicting WIfI elements from a DFU image. In terms of clinical narrative generation quality, the LLaVA-Mistral model demonstrated a 22% improvement in DFU-specific text coherence compared to the baseline model as measured by the dependency parse tree depth method. This research lays the groundwork for AI-assisted DFU assessment by creating publicly available annotations for DFU vision-language model developments and demonstrating the potential of fine-tuning to enhance clinical communication through improved classification accuracy.