Multi-scale Hybrid Transformer Network with Grouped Convolutional Embedding for Automatic Cephalometric Landmark Detection
摘要
Detection of anatomical landmarks in lateral cephalometric images is critical for orthodontic and orthognathic surgery. However, the industry faces the challenge of developing automatic cephalometric detection methods that are both precise and cost-effective for detecting as many landmarks as possible. Although current deep learning-based approaches have attained high accuracy, they have limitations in detecting landmarks that lack distinct texture features, such as certain soft tissue landmarks, and identify fewer landmarks overall. To address these limitations, we propose a novel multi-scale deep learning network that can simultaneously detect more landmarks with high accuracy, optimize model size and performance, and improve accuracy in identifying soft tissue landmarks. Firstly, we exploit a hybrid encoder that combines CNN and Swin Transformer, extracting features from different scales of the input image and fuse them. Additionally, we group 1D convolutional layers for efficient feature embedding, reducing model parameters while maintaining model features. Finally, our method achieves very high accuracy and efficiency on both public and private datasets, particularly in detecting more soft tissue landmarks with less distinct texture features.