Gastrointestinal diseases have become a significant global health issue, with early diagnosis and timely intervention being crucial for improving human health. Although endoscopy is considered the primary examination tool, it is time-consuming and invasive, limiting its widespread implementation. According to Traditional Chinese Medicine (TCM) theories, the tongue reflects the body's physiological and pathological conditions. Therefore, tongue images can be used to develop a non-invasive and rapid diagnostic model for gastrointestinal diseases. We collected tongue images from 1123 patients using a smartphone and categorized them into normal and abnormal groups based on their diagnostic reports. Comparison experiments were conducted using four models: ResNet-50, RepVGG,Deit-B and CrossViT. To address dataset imbalances, we introduced a hybrid loss function to enhance the models’ learning for underrepresented classes. The experimental results indicate that Vision Transformer demonstrates superior performance in early diagnosis of gastrointestinal diseases. Specifically, the dualbranch cross-attention model with a hybrid loss function achieved an AUC of 0.87 and an accuracy of 0.84.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Diagnosis Model of Gastrointestinal Diseases Based on Tongue Images

  • Baochen Fu,
  • Miao Duan,
  • Zhen Li,
  • Xiuli Zuo,
  • Xu Qiao

摘要

Gastrointestinal diseases have become a significant global health issue, with early diagnosis and timely intervention being crucial for improving human health. Although endoscopy is considered the primary examination tool, it is time-consuming and invasive, limiting its widespread implementation. According to Traditional Chinese Medicine (TCM) theories, the tongue reflects the body's physiological and pathological conditions. Therefore, tongue images can be used to develop a non-invasive and rapid diagnostic model for gastrointestinal diseases. We collected tongue images from 1123 patients using a smartphone and categorized them into normal and abnormal groups based on their diagnostic reports. Comparison experiments were conducted using four models: ResNet-50, RepVGG,Deit-B and CrossViT. To address dataset imbalances, we introduced a hybrid loss function to enhance the models’ learning for underrepresented classes. The experimental results indicate that Vision Transformer demonstrates superior performance in early diagnosis of gastrointestinal diseases. Specifically, the dualbranch cross-attention model with a hybrid loss function achieved an AUC of 0.87 and an accuracy of 0.84.