Tongue diagnosis plays a critical role in Traditional Chinese Medicine by providing vital clues about internal health through the analysis of tongue coating characteristics. However, the subtle color variations, low-texture regions, and heterogeneous coating patterns inherent in clinical tongue images pose significant challenges to conventional object detection frameworks. In this study, a novel deep learning approach is proposed to address these challenges by integrating advanced attention mechanisms with multi-scale feature fusion within the YOLOv8 detection framework. Specifically, the combination of Convolutional Block Attention Module (CBAM) and Feature Pyramid Net-work (FPN) is employed to enhance the extraction of both spatial-channel dependencies and hierarchical contextual features. Moreover, a targeted data augmentation strategy is introduced to enrich the training dataset, thereby im-proving the model’s ability to capture fine-grained details, especially in challenging classes such as Mirror-Approximated coatings with their near-absent, smooth appearance. Extensive evaluations from both macro- and micro-level perspectives demonstrate that the proposed enhancements significantly im-prove recall and moderate IoU detection performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing YOLOv8 for Multi-class Tongue Coating Detection: An Analysis of Attention and Multi-scale Fusion Mechanisms

  • Jiayin Li,
  • Abdur Rashid Sangi,
  • Tingfang Sun,
  • Chenyi Li

摘要

Tongue diagnosis plays a critical role in Traditional Chinese Medicine by providing vital clues about internal health through the analysis of tongue coating characteristics. However, the subtle color variations, low-texture regions, and heterogeneous coating patterns inherent in clinical tongue images pose significant challenges to conventional object detection frameworks. In this study, a novel deep learning approach is proposed to address these challenges by integrating advanced attention mechanisms with multi-scale feature fusion within the YOLOv8 detection framework. Specifically, the combination of Convolutional Block Attention Module (CBAM) and Feature Pyramid Net-work (FPN) is employed to enhance the extraction of both spatial-channel dependencies and hierarchical contextual features. Moreover, a targeted data augmentation strategy is introduced to enrich the training dataset, thereby im-proving the model’s ability to capture fine-grained details, especially in challenging classes such as Mirror-Approximated coatings with their near-absent, smooth appearance. Extensive evaluations from both macro- and micro-level perspectives demonstrate that the proposed enhancements significantly im-prove recall and moderate IoU detection performance.