Performance of a novel multimodal large language model in ınterpreting meibomian glands quantitatively and qualitatively
摘要
To evaluate the performance of a multimodal large language model (LLM), Claude 3.5 Sonnet, in interpreting meibography images for Meibomian gland dropout grading and morphological abnormality detection.
MethodsA total of 228 meibography images were analyzed by the same researcher and an assessment was performed in terms of gland drop out ratio and morphological abnormalities. Meibomian gland loss was graded from 0 (no loss) to 3 (> 2/3 loss of total gland area). One-hundred and sixty images, comprising 40 images per grade, were included. Claude 3.5 Sonnet, a multimodel LLM, developed by Anthropic (California, United States) was utilized to investigate its performance in evaluating meibography images.
ResultsClaude 3.5 Sonnet showed high performance in grading Meibomian gland dropout, correctly scoring 97.5%, 92.5%, 95%, and 85% of images in Grades 0, 1, 2, and 3, respectively. In addition, Claude 3.5 Sonnet showed remarkable performance in detecting morphological abnormalities, including heterogeneous lumen diameters, lumen tortuosity, shortened lumen length, and hyperreflective gland residues. The model detected all of the 48 manually identified morphological abnormalities accurately. In 12 images, initially classified as morphologically normal by the manual assessment, the model reported additional subtle abnormalities.
ConclusionClaude 3.5 Sonnet showed promising results in interpreting meibography images, detecting morphological abnormalities and discriminating normal Meibomian glands from abnormal. Claude 3.5 Sonnet might be useful in serving as a complementary educational tool in ophthalmology clinics. The model's ability to perform detailed morphological evaluations and respond to further questions provides a tailored learning experience for young ophthalmic clinicians.