Purpose <p>Nasopharyngeal endoscopy plays a pivotal role in the diagnosis of nasopharyngeal disorders, and the automated recognition of anatomical sites within nasopharyngeal endoscopic images is vital for enhancing clinical diagnostic accuracy, improving procedural precision, and minimizing dependence on manual interpretation.</p> Methods <p>A deep learning framework is presented for the automated recognition of anatomical sites in nasopharyngeal endoscopic images, integrating a squeeze-and-excitation (SE) block to enhance inter-class similarity and mitigate intra-class variability through refined feature extraction. To address the challenge of class imbalance during training, we employ adaptive focal loss, which dynamically prioritizes harder-to-classify samples, thereby enhancing model robustness and performance on subtle or ambiguous anatomical features.</p> Results <p>The proposed methodology was evaluated using a custom dataset consisting of 4,000 nasopharyngeal endoscopic images annotated with 20 distinct anatomical sites across six regions: nasal, nasopharynx, oropharynx, hypopharynx, laryngeal, and oral cavity areas. Among the five mainstream backbone architectures evaluated, ResNet-50 demonstrated the best performance, achieving state-of-the-art metrics with accuracy, recall, precision, and F1-score of 98.62%, 98.68%, 98.62%, and 98.63%, respectively. These results underscore the capability of our method to effectively recognize complex anatomical sites and suggest its potential for enhancing diagnostic and procedural support in the clinical setting.</p> Conclusion <p>The proposed deep learning framework addresses the challenges of inter-class similarity, intra-class variability, and class imbalance in anatomical site recognition in nasopharyngeal endoscopic images. Using ResNet-50 as the backbone, over 98% accuracy, recall, precision, and F1-score were achieved on a dataset of 4,000 images covering 20 anatomical sites. These results demonstrate the robustness of the model and its potential to enhance diagnostic accuracy and clinical workflow efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Deep Learning Method for Automated Site Recognition of Nasopharyngeal Endoscopic Images

  • Jiayin Lei,
  • Wei Yang,
  • Rongqian Yang

摘要

Purpose

Nasopharyngeal endoscopy plays a pivotal role in the diagnosis of nasopharyngeal disorders, and the automated recognition of anatomical sites within nasopharyngeal endoscopic images is vital for enhancing clinical diagnostic accuracy, improving procedural precision, and minimizing dependence on manual interpretation.

Methods

A deep learning framework is presented for the automated recognition of anatomical sites in nasopharyngeal endoscopic images, integrating a squeeze-and-excitation (SE) block to enhance inter-class similarity and mitigate intra-class variability through refined feature extraction. To address the challenge of class imbalance during training, we employ adaptive focal loss, which dynamically prioritizes harder-to-classify samples, thereby enhancing model robustness and performance on subtle or ambiguous anatomical features.

Results

The proposed methodology was evaluated using a custom dataset consisting of 4,000 nasopharyngeal endoscopic images annotated with 20 distinct anatomical sites across six regions: nasal, nasopharynx, oropharynx, hypopharynx, laryngeal, and oral cavity areas. Among the five mainstream backbone architectures evaluated, ResNet-50 demonstrated the best performance, achieving state-of-the-art metrics with accuracy, recall, precision, and F1-score of 98.62%, 98.68%, 98.62%, and 98.63%, respectively. These results underscore the capability of our method to effectively recognize complex anatomical sites and suggest its potential for enhancing diagnostic and procedural support in the clinical setting.

Conclusion

The proposed deep learning framework addresses the challenges of inter-class similarity, intra-class variability, and class imbalance in anatomical site recognition in nasopharyngeal endoscopic images. Using ResNet-50 as the backbone, over 98% accuracy, recall, precision, and F1-score were achieved on a dataset of 4,000 images covering 20 anatomical sites. These results demonstrate the robustness of the model and its potential to enhance diagnostic accuracy and clinical workflow efficiency.