Tongue diagnosis is an important diagnostic method in Traditional Chinese Medicine (TCM). Currently, while intelligent tongue diagnosis in TCM is gaining popularity, there are also some challenges, including the imbalanced amount of tongue attributes, capturing pixel-level semantic information from tongue images, and extracting discriminative information from image features and attribute dependencies. To overcome the problems mentioned above, we propose a Tongue Multi-Attribute Classification-Dual Attention Network (TMAC-DAN) consisting of three modules. The embedding extraction module optimizes TResNet by incorporating pyramid convolution for pixel-wise image feature extraction and generates initial attribute embeddings. The dual attention transformer module is designed to utilize self-attention and cross-attention mechanism to update the attribute embeddings by itself and the image features, which are subsequently fed into the attribute classifier module for attribute classification. The Tongue-attribute Group Mask Training (T-aGMT) strategy is a key component of our methodology. During training, it groups the primary attributes of the tongue and predicts masked tokens for each group based on visual features, conditioned on a set of unmasked attributes. Moreover, we constructed a new tongue image dataset named TMA including 500 tongue images for training and testing. We divide tongue features into five primary attributes and seventeen fine-grained attributes. TMAC-DAN is validated on two public tongue datasets, BioHit and PolyU/HIT, and the TMA, outperforming other state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Hybrid CNN-Transformer Network for Fine-Grained Tongue Multi-Attribute Classification with Tongue Group Attribute Mask Training in Traditional Chinese Medicine Diagnosis

  • Jiawen Liu,
  • Feixiang Ge,
  • Fufeng Li,
  • Xinlei Li,
  • Linqiang Hu,
  • Yan Rong,
  • Rongrong Zhu

摘要

Tongue diagnosis is an important diagnostic method in Traditional Chinese Medicine (TCM). Currently, while intelligent tongue diagnosis in TCM is gaining popularity, there are also some challenges, including the imbalanced amount of tongue attributes, capturing pixel-level semantic information from tongue images, and extracting discriminative information from image features and attribute dependencies. To overcome the problems mentioned above, we propose a Tongue Multi-Attribute Classification-Dual Attention Network (TMAC-DAN) consisting of three modules. The embedding extraction module optimizes TResNet by incorporating pyramid convolution for pixel-wise image feature extraction and generates initial attribute embeddings. The dual attention transformer module is designed to utilize self-attention and cross-attention mechanism to update the attribute embeddings by itself and the image features, which are subsequently fed into the attribute classifier module for attribute classification. The Tongue-attribute Group Mask Training (T-aGMT) strategy is a key component of our methodology. During training, it groups the primary attributes of the tongue and predicts masked tokens for each group based on visual features, conditioned on a set of unmasked attributes. Moreover, we constructed a new tongue image dataset named TMA including 500 tongue images for training and testing. We divide tongue features into five primary attributes and seventeen fine-grained attributes. TMAC-DAN is validated on two public tongue datasets, BioHit and PolyU/HIT, and the TMA, outperforming other state-of-the-art methods.